СБЕР

Руководитель направления Online RL (STEM) / Post-Training LLM в команду GigaChat

Мы развиваем GigaChat и ищем сильного руководителя направления online RL в домене STEM (математика, естественные науки, инженерные и технические дисциплины). Это роль для человека, который умеет одновременно развивать методы обучения модел…

СБЕР · На сервисе с: 25.09.26 17:02

Зарплата не указанаРоссияг Москва

Мы развиваем GigaChat и ищем сильного руководителя направления online RL в домене STEM (математика, естественные науки, инженерные и технические дисциплины). Это роль для человека, который умеет одновременно развивать методы обучения моделей, глубоко разбираться в предметной области и выстраивать процессы сбора и подготовки данных.

Нам нужен не просто менеджер, а сильный технический руководитель, который способен глубоко погружаться в детали, самостоятельно собирать ключевые части решения и доводить идеи до реального роста качества модели.

  • Развивать направление online RL для STEM-задач

  • Определять, как должно развиваться направление online RL в STEM-домене: какие задачи для нас наиболее важны, как измерять прогресс и что в первую очередь ограничивает рост качества модели.

  • Вести направление целиком: от постановки гипотез и плана работ до внедрения результатов в регулярный цикл обучения модели.

  • Принимать решения о приоритетах между развитием методов, сбором данных, инфраструктурой и системой оценки качества.

  • Разрабатывать и улучшать методы обучения

  • Развивать подходы post-training и online RL для задач по математике, физике, химии, биологии и другим STEM-дисциплинам.

  • Продумывать и внедрять способы оценки качества, которые помогают модели лучше решать реальные задачи: строить цепочки рассуждений, находить верные ответы, корректно применять формулы и методы, работать с многошаговыми задачами.

  • Определять, в каких случаях online RL действительно даёт прирост качества по сравнению с supervised fine-tuning и другими подходами, а в каких — нет.

  • Проводить эксперименты и разбирать результаты не только на уровне метрик, но и на уровне причин: почему модель стала лучше или хуже, насколько устойчив результат и можно ли его перенести на другие типы задач.

  • Писать ключевой код и развивать инфраструктуру

  • Самостоятельно писать и дорабатывать критичные части пайплайнов online RL.

  • Делать надёжные и воспроизводимые эксперименты: с понятными версиями данных, конфигами, сравнением запусков и контролем деградаций.

  • Выстраивать связку между моделью, верификаторами, reward-сигналами и обучающими пайплайнами так, чтобы новые идеи можно было быстро проверять и быстро доводить до практического результата.

  • Оставаться сильным инженером и исследователем, а не только руководителем: при необходимости самому разбирать узкие места в коде, экспериментах и качестве данных.

  • Строить контур данных для обучения

  • Организовывать сбор и подготовку данных для online RL в STEM-домене: задачи разной сложности, эталонные решения, формальные и автоматические верификаторы, синтетические и реальные сценарии.

  • Формировать качественные обучающие выборки с хорошим покрытием по дисциплинам, уровням сложности (от школьных до олимпиадных и университетских задач), типам рассуждений и типовым ошибкам модели.

  • Встраивать в пайплайны проверки качества: символьную и численную верификацию ответов, проверку промежуточных шагов рассуждений, контроль утечек, удаление дублей, балансировку по сложности и предметным областям.

  • Делать так, чтобы каждый цикл обучения улучшал не только модель, но и сам процесс: появлялись новые данные, новые сложные примеры, более точные критерии качества и лучшее понимание слабых мест модели.

  • Руководить сильной технической командой

  • Руководить командой исследователей и инженеров, задавать высокую планку по качеству решений, скорости работы и глубине проработки.

  • Помогать команде превращать исследовательские идеи в работающие решения, которые можно встроить в основной цикл обучения.

  • Удерживать баланс между глубиной исследований, инженерной надёжностью и практическим результатом для модели.

  • Отличное владение Python и PyTorch.

  • Практический опыт в LLM post-training: RLHF, online RL или смежных направлениях.

  • Понимание специфики STEM-домена: формальная верификация ответов, chain-of-thought reasoning, работа с математической нотацией, многошаговые решения, типовые ошибки моделей в рассуждениях.

  • Умение ставить гипотезы, проектировать эксперименты и принимать решения на основе результатов.

  • Опыт руководства сильной технической командой.

  • Готовность лично писать важные части системы руками.

Будет плюсом

  • Сильный математический или естественнонаучный бэкграунд (профильное образование, олимпиадный опыт, публикации).

  • Опыт построения верификаторов и reward-моделей для задач STEM.

  • Опыт построения пайплайнов данных, а не только работы с уже готовыми датасетами.

  • Опыт работы с distributed training или large-scale inference.

  • Опыт разработки систем оценки качества для LLM (бенчмарки, LLM-as-a-judge, process reward models).

  • Опыт работы с synthetic data generation, curriculum learning, active data collection.

  • Понимание современных open-source стеков для обучения и инференса больших языковых моделей.

  • Публикации, open-source вклад или сильный прикладной research track record.

  • Сильные и сложные задачи на переднем крае развития русскоязычных LLM.

  • Большую степень влияния на архитектуру решений, методы обучения и качество итоговой модели.

  • Команду сильных инженеров и исследователей.

  • Возможность совмещать управление направлением с глубокой технической работой.

  • Конкурентную компенсацию, премии и расширенный соцпакет.

Похожие вакансии Data Science & ML

Сайты компаний
Y

ML Developer for Speech Recognition Team (Kyrgyzstan)

yangoНа сервисе с: 03.10.26 22:14
Зарплата не указанаНе указана странаЛокация не указана
Yandex Kyrgyzstan is looking for an ML Developer to join our Speech Recognition Team. In this role, you will help improve speech recognition technologies, work with large ASR models, and contribute to launching ML solutions into production. Our team develops and improves speech recognition systems for various Yandex services, including voice assistant technologies. We are expanding our technologies to new languages and continue improving voice understanding solutions for international markets.
Yandex Kyrgyzstan is part of Yango Group, a global technology ecosystem building digital services that simplify and enhance everyday life. Operating in more than 30 countries across Europe, Africa, the Middle East, South Asia, and Latin America, we deliver solutions in ride-hailing, delivery, e-commerce, mapping, entertainment, AI, and beyond. Headquartered in Dubai, Yango Group unites advanced technological capabilities with deep local market knowledge to create products that fit naturally into people’s daily routines across diverse communities.
You will be responsible for
• Expanding speech recognition technologies • Studying and applying scientific publications, building hypotheses, and conducting experiments • Collecting and analyzing data for low-resource languages • Training large ASR models • Launching ML solutions into production • Collaborating with the team on improving speech recognition technologies
You might be a fit if you have
• Strong Python skills and experience working with PyTorch • Experience in deep learning (for example, in CV or NLP) • Basic C++ skills It’d be a plus if you: • Have worked with deep learning in speech technologies • Have experience deploying neural network models into production • Have experience working with large datasets
Explore our benefits
Health & WellbeingWork Environmentbonuses & perks
  • Life in balance with Yango

    We know how to work hard, but we also know how to have fun together, and we are true believers that the wellbeing of our teams is important
  • Psychological support

    Available in different languages to support you when you need it most
  • Life and private health cover

    Life and private health insurance
*different conditions apply based on hiring location and position type
Сайты компаний
Y

ML Developer (Kyrgyzstan)

yangoНа сервисе с: 03.10.26 22:08
Зарплата не указанаНе указана странаЛокация не указанаГибрид
Yandex Kyrgyzstan is seeking an ML developer for our LLM team building language models for Alice, with a focus on prompt interpretation, VA personality, and speech recognition/synthesis tech.
We are a small but savvy team developing a bilingual (Arabic/English) language model that enables users to seamlessly switch between languages while interacting with our voice assistant, helping them with their day-to-day needs: putting on music, setting alarms and reminders, checking the weather, controlling smart home devices, or looking something up. We then leverage this experience in our new projects focused on LLM data generation and automated QA. This is a hybrid position where you'll oversee the project's LLM core, from data preparation and model training to evaluation and deployment. The role spans the entire production cycle, giving you the opportunity to influence technical decisions from the outset. Yandex Kyrgyzstan is part of Yango Group, a global technology ecosystem building digital services that simplify and enhance daily life. Operating in more than 30 countries across Europe, Africa, the Middle East, South Asia, and Latin America, we deliver solutions in ride-hailing, delivery, e-commerce, mapping, entertainment, AI, and beyond. Headquartered in Dubai, Yango Group unites advanced technological capabilities with deep local market knowledge to create products that fit naturally into people's daily routines across diverse communities.
You will be responsible for
• Fine-tuning language models (SFT, DPO, Reinforcement Learning) • Configuring system prompts and influencing the language model's behavior • Generating prompts and multi-step conversations, improving our existing corpora, and filtering examples using language models • Comparing automated scoring (LLM-as-a-Judge) with expert annotation • Teaching the model to use the correct functions given the context, ask clarifying questions, and correctly interpret numbers, dates, names, and locations • Developing benchmarks for conversations and commands • Ensuring that the language model functions correctly as a whole (language authenticity, context memory, personality, AI safety) • Integrating models into the product infrastructure, optimizing response times and inference • Contributing to the development process from closed beta to post-launch
You might be a fit if you have
• Good knowledge of Python, with experience in developing ML/NLP systems and working with PyTorch • A solid idea of how transformers and LLMs work under the hood • Hands-on experience with fine-tuning language models using SFT, DPO, or RL (production, research, or a compelling pet project) • An understanding of QA processes: creating test datasets, analyzing data slices, using LLM judges, and running A/B tests • An independent mindset and the willingness to understand a complex system and drive results, keeping in close contact with editors and adjacent teams It’s a plus if you • Have experience with multi-language models, synthetic data, and enhancing models by improving training datasets • Know about tool calling, agent architectures, and RAG • Are familiar with LLM inference (vLLM and alternatives, distributed training) • Speak English at a level sufficient for understanding technical literature and articles
Explore our benefits
Health & WellbeingWork Environmentbonuses & perks
  • Life in balance with Yango

    We know how to work hard, but we also know how to have fun together, and we are true believers that the wellbeing of our teams is important
  • Psychological support

    Available in different languages to support you when you need it most
  • Life and private health cover

    Life and private health insurance
*different conditions apply based on hiring location and position type
Сайты компаний
T

Engineering Manager (ML)

tabbyНа сервисе с: 03.10.26 18:48
Зарплата не указанаПольшаИспанияПортугалияСербияBelgradeУдалёнка

About the role

Tabby creates financial freedom in the way people shop, earn and save by reshaping their relationship with money. Over 25 million users choose Tabby to stay in control of their spending and make the most out of their money.

The company’s flagship offering allows shoppers to split their payments online and in-store with no interest or fees. Over 70,000 global brands and small businesses, including Amazon, Noon, IKEA, and SHEIN use Tabby to accelerate growth and gain loyal customers by offering easy and flexible payments online and in stores.
Tabby generates over $18 billion in annual transaction volume for its partner brands and is the highest-rated, most-reviewed, largest, and fastest-growing FinTech in the GCC region.

Tabby launched in 2019 and has since raised +$1 billion in equity and debt funding from global and regional investors, and is now valued at $6,5 billion.

About the team

Tabby Marketplace is where our users discover what to buy. The Content Quality & Personalisation team owns the data that makes the marketplace work: a catalogue of 25M+ products from thousands of merchants, ingested through feeds and e-commerce plugins (Shopify, Salla, Zid, Amazon and more), then categorised, enriched, translated, moderated and published, largely by ML.


You will lead a cross-functional team of ML engineers, backend and frontend engineers, QA and a product analyst. The team runs the LLM-based enrichment pipeline (categorisation, attribute extraction, translation), the item representation model and embeddings that power search and recommendations, ML-assisted moderation that is replacing manual review, and the labeling and evaluation platform behind all of it.


You will work closely with the Shopping, Offers and Monetisation teams, as well as catalogue operations and partner support.

Responsibilities

  • 6+ years of engineering experience, including 3+ years building production ML systems (NLP, LLM applications, embeddings, or classification at scale)
  • 2+ years as an Engineering Manager or ML Team Lead at a fast-growing e-commerce, marketplace or fintech company
  • Hands-on experience shipping LLM-based products: prompt and pipeline design, fine-tuning, evaluation, cost and latency control, self-hosted and API-based models
  • Experience building and operating large-scale data and ML pipelines (batch and streaming), and making them observable, reproducible and reliable
  • Solid backend fundamentals; you are comfortable reviewing Go and Python services and reasoning about distributed systems
  • Our stack: Python, Go, PostgreSQL, Pub/Sub, BigQuery, GCS, Kubernetes, Google Cloud Platform, Airflow, and a microservices architecture
  • A strong grasp of ML evaluation: golden datasets, labeling workflows, offline metrics, and A/B testing tied to business outcomes
  • Product sense: you connect catalogue quality to conversion, discovery and merchant growth, and you can prioritise accordingly
  • A proactive mindset and the ability to work independently
  • Strong communication skills in English (B2 level or higher)

Nice to have:
  • Experience with product catalogues, PIM systems, or marketplace content moderation
  • Experience with Arabic-language content
  • Familiarity with data residency and regulated-data requirements

Qualifications

  • Own the end-to-end product data pipeline: ingestion from feeds and plugins, ML enrichment, moderation and publication, with clear SLAs for freshness, coverage and quality
  • Lead the ML roadmap for catalogue intelligence: category tree and attribute coverage, translation quality, ML-assisted moderation, item embeddings and recommendations
  • Lead large cross-team projects and drive them to production
  • Contribute to quarterly planning and roadmap definition; define and report OKRs for catalogue quality and personalisation
  • Review feature designs and ensure non-functional requirements are met, including ML evaluation, inference cost, latency and data residency
  • Build and maintain the evaluation and labeling infrastructure that lets the team measure every model change before it reaches production
  • Oversee technical debt management and incident handling across ML and backend services
  • Hire, evaluate, and motivate team members; grow ML engineers into owners of business outcomes
  • Build cross-team and cross-functional collaboration with Shopping, Offers, Monetisation, catalogue operations and partner support to increase efficiency
  • Foster a results- and business-oriented culture
  • Monitor key team performance indicators
  • Ensure process and delivery transparency for stakeholders and partner functions
  • Optimise processes to improve productivity

Benefits

  • Full-time B2B contract
  • Fully remote setup
  • Up to 20% tax allowance
  • 22 paid leave days annually
  • Stock options (ESOP) in a fast-scaling, pre-IPO company
  • Flexi benefits you can use for wellness, travel, or learning
  • Work alongside a high-performing, international engineering team in a global fintech unicorn

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.