яндекс

Разработчик на Go в Server Infrastructure Yandex Cloud (Python)

Команда Server Infrastructure занимается эксплуатацией быстро растущей инфраструктуры Yandex Cloud в рамках подразделения Cloud Foundation Services. Мы строим надёжную и масштабируемую инфраструктуру, поверх которой запускаются виртуальные…

яндекс На сервисе с: 09.10.26 08:22

Зарплата не указанаНе указана странаЛокация не указана

Команда Server Infrastructure занимается эксплуатацией быстро растущей инфраструктуры Yandex Cloud в рамках подразделения Cloud Foundation Services. Мы строим надёжную и масштабируемую инфраструктуру, поверх которой запускаются виртуальные машины пользователей и внутренние сервисы. В сервисах реализуем различные сценарии работы с железом: от процессов ввода, вывода, починки до бесшовного обновления ОС на всём кластере.

Наши сервисы работают с большим количеством облачных и общих яндексовых систем, собирают данные о хостах, метрики состояния железа и кластера в целом, чтобы планировать обслуживание серверов и распределять ресурсы. Мы предоставляем сервисы и инструменты, которые упрощают и автоматизируют внутренние процессы, делают инфраструктуру прозрачнее и стабильнее, снимают с инженеров рутинную работу.

Под нашим управлением уже более 20 тыс. серверов в трёх дата-центрах Яндекса, и их количество непрерывно растёт. Мы разрабатываем и постоянно совершенствуем способы мониторинга наших серверов и подходы к нему так, чтобы заранее и автоматически диагностировать неполадки и выполнять обслуживание, не дожидаясь выхода серверов из строя.

В работе мы используем:
* Golang и Python для разработки сервисов и автоматики
* SaltStack и Terraform для описания инфраструктуры
* TeamCity и Spinnaker для процессов CI/CD

Автоматизация жизненного цикла серверов

Вам предстоит развивать систему эксплуатации инфраструктуры Yandex Cloud:
заменять ручные операции автоматикой — от ввода и вывода серверов
до бесшовного обновления ОС на всём кластере.

Разработка системы инвентаризации

Вы будете проектировать и дорабатывать систему инвентаризации, которая
объединяет физическую и логическую информацию обо всём парке железа, —
чтобы хранить и использовать данные об инфраструктуре было проще.

Разработка для Kubernetes

Вы будете участвовать в разработке Kubernetes-операторов и устройств на базе device plugin, а также интегрировать NRI (Node Runtime Interface) — инструменты, через которые облако управляет ресурсами узлов кластера и контейнерами
на хостах.

Разработка сервисов мониторинга

Вы будете заниматься разработкой сервисов-агентов для мониторинга, сбора метрик
и выполнения проверок на хостах — чтобы неполадки диагностировались
автоматически, до выхода сервера из строя.

Развитие инструментов планирования ресурсов

Также вы займётесь поддержкой и развитием инструментов, которые помогают планировать и распределять ресурсы внутри облака.

Больше о том, что под капотом платформы Yandex Cloud, — в канале Inside Yandex Cloud

* Пишете чистый код на Go — либо на Python и готовы быстро освоить Go
* Разрабатывали и поддерживали отказоустойчивые системы
* Выстраивали и эксплуатировали CI/CD-процессы для сервисов, использовали концепцию Infrasructure as Code
* Знаете, как построить идеальный мониторинг
* Любите улучшать процессы и автоматизировать задачи, писали сервисы и утилиты для автоматизации

* Работали с системами контейнеризации и управления конфигурациями: Docker, K8s, SaltStack, Ansible, Terraform
* Разрабатывали операторы и плагины для K8s
* Строили облачные сервисы
* Глубоко знаете Linux

Похожие вакансии Go

Сайты компаний
ozon

Старший Go-разработчик, ML Workflow

ozonНа сервисе с: 06.08.26 15:35↑ Вакансия с автоподнятием
Зарплата не указанаРоссияМоскваУдалёнка

Привет! Это команда ML Платформы. Мы отвечаем за стандартизацию процессов MLOps и ModelOps: разрабатываем и поддерживаем сервисы и пакеты для запуска пайплайнов машинного обучения, трекинга экспериментов, хранения и версионирования моделей и датасетов, а также инференса и мониторинга моделей.

Наша команда разрабатывает сервисы для всех дата-сайентистов, адаптирует инструменты ML-инфраструктуры под потребности отдельных команд, занимается кастомными интеграциями и исследованиями.

Ищем инженера, который готов вместе с нами развивать платформенную инфраструктуру и поддерживать высокий уровень инженерной культуры в компании.

Наш стек

  • Go с платформенными интеграциями, Kubernetes, GitLab CI, PostgreSQL, Redis, ClickHouse, Kafka, Python для скриптов.

Примеры задач

  • Поддержка совместной разработки для нескольких пользователей одновременно в jupyterhub.

  • Разработка сервиса по запуску интерактивных контейнеров на GPU.

  • Интеграция с платформенными сервисами по выделению общих хранилищ для команд.

Вы будете

  • Разрабатывать и развивать сервисы ML-платформы.

  • Проектировать микросервисную архитектуру и интегрироваться с платформенными сервисами.

  • Участвовать в код-ревью и поддерживать качество кодовой базы.

  • Исследовать и внедрять новые инструменты для улучшения ML-инфраструктуры.

Нам важно

  • Опыт коммерческой разработки на Go от трех лет, понимание тонкостей языка.

  • Опыт работы с Kubernetes, умение дебажить проблемы с помощью kubectl.

  • Понимание принципа «at least once».

  • Опыт проектирования микросервисной архитектуры.

  • Желание строить качественную ML-платформу и развивать её.

Будет плюсом

  • Умение читать и дебажить код на Python.

  • Опыт работы с JupyterHub.

  • Опыт построения дашбордов в Grafana.

  • Опыт работы с ML и понимание MLOps-практик.

Эта вакансия также есть на:hh.ru
Сайты компаний
N

Senior System Engineer (Virtual Private Cloud Team)

nebiusНа сервисе с: 06.10.26 01:26↑ Вакансия с автоподнятием
Зарплата не указанаНидерландыAmsterdam

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

The VPC (Virtual Private Cloud) team builds and operates the core networking layer of our cloud platform. We enable virtual machines to communicate reliably by abstracting them from the underlying physical network, providing seamless IP-based connectivity even as workloads move across the datacenter.

Our team designs secure, isolated private networks for customers, connects them to the public internet, and delivers essential networking services such as load balancers, firewalls, and security groups. Because nearly every cloud application depends on networking, VPC is a foundational service for the entire platform.

Stability and resilience are our top priorities—we engineer systems to remain reliable and operational even when parts of the cloud experience failures.

Our team both develops the control-plane, written in Golang, and supports open-source based dataplane, written mostly in C. This makes us quite unique in that we both have high-level control plane tasks and low-level system research tasks, with different team members focusing more on one part or the other.

What we do:

  • Increase service scalability and reliability
  • Develop the control & data plane for Nebius Cloud network services
  • Develop service functionality (e.g. IPv6 support, service endpoints, VPC peering, L7 load balancers and other networking services)
  • Improve the internal architecture, optimizing interaction with related services
  • Debug datapath and kernel issues, including performance related issues
  • Perform load testing for services

We expect you to have:

  • Capable of writing reliable, high-performance concurrent code
  • Proficient in Go and able to read C/C++,  or be ready to learn both
  • Familiar with routing network traffic

It would be an added bonus if you:

  • Worked with OVS, VPP, DPDK, Linux kernel network subsystem or other network dataplanes
  • Well-versed in virtual networks and overlays, SDN, NFV, DPI, network protocols, routing and tunneling, BGP, and MPL

We conduct coding interviews as part of the process.

#LI-JS1

Benefits & Perks:

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

What's it like to work at Nebius:

Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI 

Equal Opportunity Statement:

Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.

Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. 

If you need accommodations during the application process, please let us know.

Сайты компаний
P

Engineering Manager - Golang [Telesales]

plataНа сервисе с: 03.10.26 17:45↑ Вакансия с автоподнятием
Зарплата не указанаВесь мир

We are looking for an Engineering manager to lead our Telesales team.

This team is engaged in tasks related to customer acquisition via phone sales. The team enables work of a number of call centers, providing working web environment and telephony capabilities to sales agents. Telesales is responsible for web-tools for client acquisition, lead call strategies, agent performance evaluation and more. It is also planned to develop a platform based on the team’s solution, so you will have the opportunity to directly influence the end-result. 

This role is primarily about leading and growing people, sustaining and sharpening a team that already works well as well as being proactive with the business. In practice, that means you act as a manager and mentor for the team – developing engineers, shaping processes, and ensuring delivery quality, while still staying close enough to the codebase to make informed technical decisions and occasionally contribute to implementation. 

 

Challenges that await you:

  • Lead the team: manage delivery, prioritize new features demanded by business, while contributing to development as needed
  • Develop talents: provide feedback, enable growth for colleagues, support the team, and mentor
  • Overseeing team performance, ensuring on-time project delivery aligned with business goals
  • Solving complex challenges and ensuring high-quality solutions
  • Shaping the architecture, design, and implementation of backend services in Golang

What makes you a great fit:

  • 1+ years of experience managing engineering teams in fast-paced settings, preferably fintech with Go stack
  • Hiring, development, performance, and delivery management and strong skills in cross-functional collaboration and the ability to build partnerships
  • A strong technical background and development experience, enabling you to participate in architectural discussions and validate the team's technical solutions
  • 4+ years of experience as a Go developer
  • Comfort operating with loosely defined business requirements — able to set priorities independently and defend them with clear reasoning 
  • B1 or higher English level for effective communication with an international team

Our ways of working:

  • Innovative Spirit: A commitment to creativity and groundbreaking solutions
  • Honest Feedback: valuing open, transparent communication
  • Supportive Team: a strong, collaborative community
  • Celebrating Achievements: recognizing our wins together
  • High-Tech Environment: a team full of smart and revolutionary people who date to challenge the status quo of incumbent finances

Our benefits:

  • Relocation support to one of our hubs — Cyprus, Serbia, Georgia, Spain— with assistance for the employee and their family
  • Flexible work from one of our offices or remote
  • Healthcare Coverage
  • Education Budget: Language lessons, professional training and certifications
  • Wellness Budget: Mental health and fitness activity reimbursements
  • Vacation policy: 20 days of annual leave and paid sick leave

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.