СБЕР

Senior Research Engineer (LLM Pretraining)

Мы занимаемся pretrain'ом больших языковых моделей в GigaChat: проектируем архитектуру, подбираем рецепт обучения и поддерживаем весь инженерный контур вокруг него.

СБЕР · На сервисе с: 25.09.26 18:03

Зарплата не указанаРоссияг Москва

Мы занимаемся pretrain'ом больших языковых моделей в GigaChat: проектируем архитектуру, подбираем рецепт обучения и поддерживаем весь инженерный контур вокруг него.

Недавно мы обучили MoE-модель на 700 миллиардов параметров — и на этом не собираемся останавливаться. Обучение идёт на кластерах H100 и B200. GigaChat — самый быстрорастущий проект Сбера, и pretrain — его ядро.

Чем занимается команда

  • Архитектура и законы масштабирования.
  • Рецепт обучения: оптимизаторы, расписание learning rate, нормализация, точность вычислений.
  • Устойчивость больших прогонов и ускорение сходимости.
  • Диагностика обучения и оценка изменений с опорой на математический аппарат.
  • Инженерный контур: воспроизводимость, тесты, CI/CD.
  • Роль с акцентом на модель, оптимизацию и инфраструктуру обучения, а не на данные. Главная цель — делать обучение быстрее, надёжнее и предсказуемее.

Какие задачи стоят перед командой

  • Ускорение цикла «идея → эксперимент → вывод → внедрение».
  • Снижение количества ручных прогонов и неочевидных сбоев, повышение воспроизводимости и прозрачности результатов.
  • Повышение надёжности больших прогонов.
  • Ранняя диагностика деградаций и отделение реальных улучшений от ложных сигналов (расхождение, NaN, коллапс энтропии, артефакты маршрутизации, ложное снижение функции потерь).
  • Обеспечение безопасного масштабирования при внедрении крупных архитектурных изменений.
  • Анализ влияния сложных архитектур (например, mixture of experts и маршрутизация) на качество, стабильность и скорость обучения.
  • Определение и развитие метрик, корректно отражающих изменения в обучении моделей.

Почему мы

Масштаб - 700B MoE уже обучена, дальше — больше. Кластеры на H100 и B200.

Публикации. Можно и нужно писать статьи по результатам своей работы — это не ограничивается.

Команда - в России нет другой команды, которая занимается pretrain'ом на таком масштабе. Коллеги — люди, которые глубоко разбираются в теме.

Влияние. Вы берёте направление целиком. Это не «выполнять задачи из бэклога», а самостоятельно определять, что важно, и доводить до результата.

  • Взять на себя целое направление внутри pretrain'а и развивать его: от постановки задач и планирования экспериментов до внедрения результатов в основное обучение.

  • Проектировать и проводить эксперименты: формулировать гипотезы, запускать абляции, сравнивать подходы, разбираться в результатах и превращать выводы в решения для основного обучения.

  • Разбираться с нестабильностью на больших прогонах: искать причины деградаций, строить диагностические метрики, предлагать изменения в оптимизаторе, расписании lr, нормализациях, инициализации, клиппинге, точности вычислений и маршрутизации.

  • Работать с архитектурой смеси экспертов (MoE): маршрутизатор, балансировка нагрузки, переполнение, артефакты маршрутизации, влияние на качество и производительность.

  • Поддерживать большие прогоны и продолжения обучения с чекпоинтов: следить за дрейфом, проверять изменения в коде и конфигурации, снижать риск регрессий.

  • Улучшать инженерное качество контура обучения: ревью критичных изменений, стратегия тестирования, воспроизводимость экспериментов, профилирование и устранение узких мест.

  • Глубокое понимание устройства обучения нейросетей: не на уровне обзоров и пересказов, а на уровне, где вы можете объяснить, почему конкретный прогон расходится, глядя на кривые функции потерь, нормы градиентов и энтропии.

  • Способность самостоятельно взять направление и довести его до результата: от чтения статей и постановки гипотез до внедрения в основной трейн.

  • Практический опыт с PyTorch и именно с обучением моделей, а не только с инференсом.

  • Умение доводить исследовательские идеи до надёжного инженерного решения: воспроизводимость, конфиги, тесты, автоматизация, понятные критерии качества.

  • Хорошую инженерную культуру: аккуратные PR, профилирование, внимание к качеству кода, понятные отчёты об экспериментах.

Будет плюсом

  • Опыт со смешанной точностью и распределённым обучением.

  • Опыт построения систем оценки моделей или инфраструктуры для экспериментов.

  • Удалённо.

  • Возможность оформления в аккредитованную IT-компанию.

  • Годовая премия по итогам работы до 6 окладов.

  • Регулярный пересмотр зарплат.

  • Корпоративный спортзал и зоны отдыха.

  • Более 400 программ СберУниверситета для роста.

  • Программа адаптации и помощь руководителя на старте.

  • Крупнейшее DS&AI community — более 600 DS банка, регулярный обмен знаниями, опытом и лучшими практиками, интерактивные лекции и мастер-классы от ведущих ВУЗов и экспертов технологических компаний, дайджест о самых последних разработках в области DS&AI и отчеты с крупнейших конференций мира, регулярные внутренние митапы.

  • Расширенный ДМС, льготное страхование для семьи, корпоративная пенсионная программа.

  • Ипотека для сотрудников по дисконтной программе.

  • СберПрайм+ и скидки у партнёров.

  • Бонус за рекомендации в команду

Эта вакансия также есть на:hh.ru

Похожие вакансии Data Science & ML

Сайты компаний
Y

ML Developer for Speech Recognition Team (Kyrgyzstan)

yangoНа сервисе с: 03.10.26 22:14
Зарплата не указанаНе указана странаЛокация не указана
Yandex Kyrgyzstan is looking for an ML Developer to join our Speech Recognition Team. In this role, you will help improve speech recognition technologies, work with large ASR models, and contribute to launching ML solutions into production. Our team develops and improves speech recognition systems for various Yandex services, including voice assistant technologies. We are expanding our technologies to new languages and continue improving voice understanding solutions for international markets.
Yandex Kyrgyzstan is part of Yango Group, a global technology ecosystem building digital services that simplify and enhance everyday life. Operating in more than 30 countries across Europe, Africa, the Middle East, South Asia, and Latin America, we deliver solutions in ride-hailing, delivery, e-commerce, mapping, entertainment, AI, and beyond. Headquartered in Dubai, Yango Group unites advanced technological capabilities with deep local market knowledge to create products that fit naturally into people’s daily routines across diverse communities.
You will be responsible for
• Expanding speech recognition technologies • Studying and applying scientific publications, building hypotheses, and conducting experiments • Collecting and analyzing data for low-resource languages • Training large ASR models • Launching ML solutions into production • Collaborating with the team on improving speech recognition technologies
You might be a fit if you have
• Strong Python skills and experience working with PyTorch • Experience in deep learning (for example, in CV or NLP) • Basic C++ skills It’d be a plus if you: • Have worked with deep learning in speech technologies • Have experience deploying neural network models into production • Have experience working with large datasets
Explore our benefits
Health & WellbeingWork Environmentbonuses & perks
  • Life in balance with Yango

    We know how to work hard, but we also know how to have fun together, and we are true believers that the wellbeing of our teams is important
  • Psychological support

    Available in different languages to support you when you need it most
  • Life and private health cover

    Life and private health insurance
*different conditions apply based on hiring location and position type
Сайты компаний
Y

ML Developer (Kyrgyzstan)

yangoНа сервисе с: 03.10.26 22:08
Зарплата не указанаНе указана странаЛокация не указанаГибрид
Yandex Kyrgyzstan is seeking an ML developer for our LLM team building language models for Alice, with a focus on prompt interpretation, VA personality, and speech recognition/synthesis tech.
We are a small but savvy team developing a bilingual (Arabic/English) language model that enables users to seamlessly switch between languages while interacting with our voice assistant, helping them with their day-to-day needs: putting on music, setting alarms and reminders, checking the weather, controlling smart home devices, or looking something up. We then leverage this experience in our new projects focused on LLM data generation and automated QA. This is a hybrid position where you'll oversee the project's LLM core, from data preparation and model training to evaluation and deployment. The role spans the entire production cycle, giving you the opportunity to influence technical decisions from the outset. Yandex Kyrgyzstan is part of Yango Group, a global technology ecosystem building digital services that simplify and enhance daily life. Operating in more than 30 countries across Europe, Africa, the Middle East, South Asia, and Latin America, we deliver solutions in ride-hailing, delivery, e-commerce, mapping, entertainment, AI, and beyond. Headquartered in Dubai, Yango Group unites advanced technological capabilities with deep local market knowledge to create products that fit naturally into people's daily routines across diverse communities.
You will be responsible for
• Fine-tuning language models (SFT, DPO, Reinforcement Learning) • Configuring system prompts and influencing the language model's behavior • Generating prompts and multi-step conversations, improving our existing corpora, and filtering examples using language models • Comparing automated scoring (LLM-as-a-Judge) with expert annotation • Teaching the model to use the correct functions given the context, ask clarifying questions, and correctly interpret numbers, dates, names, and locations • Developing benchmarks for conversations and commands • Ensuring that the language model functions correctly as a whole (language authenticity, context memory, personality, AI safety) • Integrating models into the product infrastructure, optimizing response times and inference • Contributing to the development process from closed beta to post-launch
You might be a fit if you have
• Good knowledge of Python, with experience in developing ML/NLP systems and working with PyTorch • A solid idea of how transformers and LLMs work under the hood • Hands-on experience with fine-tuning language models using SFT, DPO, or RL (production, research, or a compelling pet project) • An understanding of QA processes: creating test datasets, analyzing data slices, using LLM judges, and running A/B tests • An independent mindset and the willingness to understand a complex system and drive results, keeping in close contact with editors and adjacent teams It’s a plus if you • Have experience with multi-language models, synthetic data, and enhancing models by improving training datasets • Know about tool calling, agent architectures, and RAG • Are familiar with LLM inference (vLLM and alternatives, distributed training) • Speak English at a level sufficient for understanding technical literature and articles
Explore our benefits
Health & WellbeingWork Environmentbonuses & perks
  • Life in balance with Yango

    We know how to work hard, but we also know how to have fun together, and we are true believers that the wellbeing of our teams is important
  • Psychological support

    Available in different languages to support you when you need it most
  • Life and private health cover

    Life and private health insurance
*different conditions apply based on hiring location and position type
Сайты компаний
V

AI/ML Intern Summer 2027

veeamНа сервисе с: 03.10.26 21:39
Зарплата не указанаКанадаSan Jose

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About Out Summer Internship Program 

Our Summer Internship Program is designed for students entering their final year of university who are eager to gain meaningful, real-world experience in a fast paced, collaborative, and professional environment. 

As a Summer Intern, you'll participate in a comprehensive onboarding experience led by our University Relations team to set you up for success from day one. Throughout the program, you'll also have the opportunity to participate in weekly professional development sessions, networking events, social activities, and other engaging experienced designed to support your personal and professional growth. 

The program takes place from June – August 2027 (10-week program).  

What We're Looking For

  • Passion & Curiosity: Strong interest in machine learning research, experimentation, and understanding model behavior.
  • Machine Learning Fundamentals: Strong foundation in supervised/unsupervised learning, optimization, regularization, model evaluation, and deep learning fundamentals.
  • Model Training Experience: Hands-on experience training deep learning models in PyTorch or TensorFlow. Ability to diagnose poor convergence, overfitting, unstable training, gradient issues, data leakage, and weak generalization.
  • Statistics & Experimentation: Strong understanding of probability, statistics, hypothesis testing, experimental analysis, and interpreting noisy results.
  • Software Engineering Discipline: Ability to write clean, maintainable code with strong encapsulation, separation of concerns, modularity, and object-oriented design principles.
  • Rapid Prototyping: Comfortable using Claude or similar AI tools for development, debugging, and rapid iteration.
  • Research Mindset: Ability to independently investigate problems, design experiments, and analyze outcomes critically.

Nice To Have

  • LLM Experience: Experience training, fine-tuning, or evaluating transformer models or LLMs.
  • Modern ML Tooling: Familiarity with Weights & Biases, MLflow, distributed training, mixed precision, LoRA/QLoRA, or hyperparameteroptimization.
  • Research Exposure: Experience reproducing papers, participating in ML competitions, contributing to research projects, or building advanced personal projects.
  • Applied AI Domains: Exposure to NLP, generative AI, multimodal systems, retrieval systems, or recommendation systems.

What You Could Be Working On

  • Model Training & Evaluation: Train and improve ML models across a variety of datasets and tasks.
  • Training Diagnostics: Analyze loss curves, gradients, metrics, and experiments to diagnose model failures and improve performance.
  • LLM & Generative AI Research: Work on transformer models, fine-tuning workflows, evaluation systems, and generative AI applications.
  • Rapid Experimentation: Prototype and iterate quickly using Claude-assisted development workflows.
  • Research Tooling: Build reusable experimentation, training, and evaluation workflows for ML research.

Candidates should have completed advanced coursework in areas such as:

  • Machine Learning
  • Deep Learning
  • Probability & Statistics
  • Linear Algebra
  • Optimization
  • Algorithms & Data Structures
  • Artificial Intelligence
  • Natural Language Processing
  • Computer Vision
  • Reinforcement Learning
  • Software Engineering

Targeted Field of Study 

Currently pursuing a Master’s degree or PhD in: Computer Science, Artificial Intelligence, Machine Learning, Statistics, Applied Mathematics, Data Science, Electrical Engineering, Or other closely related quantitative fields

Requirements:

  • This role requires you to be in office 5 days a week at the San Jose, California location

Benefits

As a paid intern at Veeam, you’ll receive:

  • Paid Company Holidays during your internship
  • Tech Stipend to help set up your workspace
  • 8 Hours of Paid Volunteer Time through our Veeam Cares Program
  • Personal and Professional Development through our Internship Program

 

We’re committed to providing a supportive and rewarding internship experience. 

The pay range posted is an hourly rate of base pay. When making an offer of employment, Veeam will take into consideration the candidate’s expectations, experience, education, scope of responsibility for the role, and the current market demands.

United States of America Intern Pay Range

$40 - $50 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.