мфк вэббанкир

Tech Lead MLOps

Мы строим современную платформу данных и ML: хранилище по слоям, единый профиль клиента (MDM), фичестор и систему принятия решений по заявкам в реальном времени.

мфк вэббанкир · На сервисе с: 28.09.26 16:20

Зарплата не указанаРоссияМоскваУдалёнка

Мы строим современную платформу данных и ML: хранилище по слоям, единый профиль клиента (MDM), фичестор и систему принятия решений по заявкам в реальном времени.

ML-инфраструктуры в полноценном виде у нас пока нет, и мы ищем человека, который спроектирует её сам: выберет подходы и инструменты, обоснует решения, реализует и станет владельцем платформы. Это не позиция «делать по готовому ТЗ».

Важно: это роль технического лидера, а не менеджера. Основную часть времени вы будете сами, своими руками разворачивать, настраивать и отлаживать инфраструктуру — от кластеров и сервисов до CI/CD, логирования и алертов.

Чем предстоит заниматься:

  • Спроектировать целевую архитектуру ML-платформы: как модели обучаются, выкатываются, обслуживаются и мониторятся; как фичи попадают из DWH в онлайн-контур скоринга.
  • Сравнивать и выбирать технологии под наши задачи, готовить обоснования и защищать решения перед командой и руководством.
  • Своими руками развернуть и настроить: dbt поверх Greenplum, Yandex Data Proc (Spark), MLflow (трекинг экспериментов, реестр моделей), serving моделей (онлайн-инференс по API и батч-скоринг, версионирование, откаты).
  • Развернуть, настроить и поддерживать инфраструктуру фичестора на Feast: офлайн-хранилище в DWH (Greenplum), онлайн-хранилище на Valkey, registry, materialization, логирование и мониторинг. Сами признаки и пайплайны строят дата-инженеры, а вы задаёте стандарты работы с фичестором (регистрация фич, point-in-time корректность, отсутствие расхождений train/serve), обучаете команду и ревьюите решения.
  • Выстроить мониторинг на всех уровнях: инфраструктура и сервисы — метрики, логи, алертинг (Prometheus, Grafana, Alertmanager, ELK/Loki); модели — дрейф фич и скоров, стабильность (PSI), деградация качества, латентность и ошибки инференса; данные — свежесть, полнота и корректность фич и витрин, контроль пайплайнов Airflow и dbt; дашборды и регламенты реагирования на инциденты.
  • Обеспечить онлайн-скоринг заявки с низкой задержкой: сбор фич из фичестора, MDM и внешних источников.
  • Автоматизировать пайплайны (Airflow, Kafka, Debezium), выстроить инфраструктуру как код (Terraform) и CI/CD.
  • Поддерживать всё развёрнутое в эксплуатации: обновления, масштабирование, разбор инцидентов, runbook-и.
  • Самостоятельно находить узкие места и риски в платформе и предлагать, что улучшить, не дожидаясь запроса.
  • Задавать стандарты MLOps/DataOps и помогать командам DWH и Data Science работать по ним.

Что мы ждём:

  • От 5 лет в DevOps / Data Engineering / MLOps, из них от 2 лет — ML-инфраструктура в продакшене.
  • Сильный hands-on опыт: вы сами разворачивали и эксплуатировали инфраструктуру в проде и готовы много работать руками — Kubernetes, Terraform, конфиги, пайплайны, отладка.
  • Опыт проектирования платформы или крупной подсистемы с нуля — не только поддержки существующей. Готовность рассказать, какие решения принимали и почему.
  • Архитектурное мышление: умение видеть систему целиком, учитывать нагрузку, отказоустойчивость, стоимость владения и безопасность, взвешивать компромиссы.
  • Опыт вывода ML-моделей в прод с онлайн-инференсом (FastAPI, BentoML, KServe, Triton или аналоги).
  • Уверенный Python, SQL, Linux, Git.
  • MLflow (или аналоги), Airflow, Spark, Kafka.
  • Docker, Kubernetes, CI/CD, Terraform.
  • dbt и MPP-хранилища (Greenplum, PostgreSQL, ClickHouse), + Redis.
  • Опыт развёртывания и эксплуатации Feast в проде: online/offline store, registry, materialization, point-in-time корректность, интеграция с сервисом инференса. Умение передать эти знания дата-инженерам: написать гайдлайны, провести ревью, разобрать ошибки.
  • Опыт построения мониторинга с нуля: Prometheus, Grafana, алертинг; мониторинг ML-моделей (Evidently, NannyML или собственные решения); контроль качества данных (dbt tests, Great Expectations или аналоги).
  • Проактивность: вы сами формулируете задачи, доводите их до результата и приходите с предложениями, а не ждёте постановки.
  • Умение письменно оформлять архитектурные решения (ADR, схемы, документация) и объяснять их техническим и нетехническим коллегам.

Будет плюсом:

  • Опыт в Yandex Cloud (Data Proc, Managed Kubernetes, Managed Greenplum / MPP Analytics).
  • Опыт в финтехе, банках или МФО: кредитный скоринг, антифрод, работа с БКИ.
  • Знание требований 152-ФЗ и практик работы с персональными данными.
  • Опыт с потоковой обработкой (Flink, Spark Structured Streaming).

Мы предлагаем:

  • Возможность построить ML-платформу с нуля и влиять на архитектуру.
  • Современный стек и реальные задачи с измеримым влиянием на бизнес.
  • Зарплата: обсуждается по итогам интервью.
  • Формат работы: удобный для вас [удалённо / гибрид / офис в Москве].
  • Скидки от компаний-партнеров через PrimeZone.

Похожие вакансии MLOps инженер

Сайты компаний
T

Senior Machine Learning Engineer

tolokaНа сервисе с: 03.10.26 19:22↑ Вакансия с автоподнятием
Зарплата не указанаЕвропаПольшаКипрВеликобританияГерманияНидерландыИспанияПортугалияСербияЧехияЛюксембургМолдоваЛатвияЧерногорияЛитваШвейцарияЭстонияВенгрияИсландияНорвегияРумынияФинляндияФранцияАвстрияГрецияДанияИталияСловакияУдалёнка

About Toloka

At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds, representing over 100 countries and speaking 40+ languages. We are a well-funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.

Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board. Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.

 

About the Team

We are the ML team inside Toloka — we build the machine-learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.

A few examples of what we own:

  • LLM QA — the core technology behind Toloka's automated quality-check mechanism. Every annotation flowing through Self-Service is reviewed by an LLM agent we design, train, and operate.
  • Model distillation and fine-tuning — adapting frontier and open-source models to Toloka's tasks to hit the right quality at the right cost.
  • Evaluation, benchmarking, cost modeling, and model selection across providers.

We own the full chain. The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration. No hand-off between research, engineering, and operations — it's all us.

 

About the Position

As a Senior ML Engineer, you will design, train, and deploy the AI agents that drive our core products. You will focus on end-to-end ML tasks—including fine-tuning and Reinforcement Learning (RL)—to build resilient agentic workflows in Python. You will own the full lifecycle of your models, closing the loop from research and benchmarking to production scaling and monitoring.

 

What you’ll do

  • Train, fine-tune, and distill ML models (including RL approaches) to power autonomous AI agents.
  • Build and operate agentic workflows in Python, handling complex reasoning and hybrid human-expert interactions.
  • Own evaluation and benchmarking, selecting foundational models and establishing cost models.
  • Manage the full ML lifecycle: design solutions, ship to production, and monitor real-time signals.
  • Implement observability metrics tailored for agent logic, model performance, and system reliability.

 

What we're looking for

  • 3+ years of experience in ML: Strong background in model training, fine-tuning, and Reinforcement Learning (RL);
  • 1+ year in Agent Development: Practical experience building, evaluating, and launching autonomous AI agents;
  • Agentic Frameworks: Proven experience working with frameworks such as LangChain, LlamaIndex, or AutoGen;
  • Open-Source LLMs: Practical knowledge of model distillation and adapting open-source models (e.g., Llama, Mistral);
  • Python Mastery: Advanced proficiency in Python with a drive to apply disciplined software engineering standards to ML;
  • End-to-End Ownership: Ability to work across the entire chain, from research to production operations;
  • Language: Fluency in English (B2 or above).

 

What we can offer

  • You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
  • Competitive compensation package including base salary, bonus, and ESOP.
  • Paid PTO and benefits will vary depending on location.
  • We offer a full remote or hybrid model (if you are based in NL or Serbia).
  • IT setup and home office allowances.

 

Equal Opportunity Employer:

Toloka is committed to providing equal opportunity and fostering an inclusive environment. We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law. Selection decisions are made based on qualifications, merit, and business need.

 

[Important Notice] Scam Alert Regarding Fake Job Postings

It has come to our attention that an individual or group is fraudulently impersonating Toloka to post fake jobs and solicit personal information from applicants. Please be aware:

  • Official Communication: Our recruiting team will only contact you from an official "toloka.ai" email address. We will NEVER use Gmail, Yahoo, Tolokainc, toloka.inc, or other personal or seemingly business email accounts.
  • Our Process: We will never ask for your bank account details, credit card number, or any fees as part of the application or interview process.
  • Official Listings: All legitimate job openings are posted on our official careers page: https://toloka.ai/careers#job-list

What to do: If you see a suspicious job posting or have been contacted by someone you suspect is a scammer, please do not provide any personal information. Instead, report the incident to us directly at security@toloka.ai and report the profile/post to LinkedIn.We are taking this matter very seriously and are working with the appropriate parties to resolve it.

Thank you for your vigilance!

To learn how we collect, use, disclose, and store personal data, check out our Privacy Notice.

Сайты компаний
J

Research Engineer (Agentic Models)

jetbrainsНа сервисе с: 03.10.26 17:14↑ Вакансия с автоподнятием
Зарплата не указанаКипрВеликобританияГерманияНидерландыИспанияСербияЧехияAmsterdam

At JetBrains, code is our passion. Ever since we started, back in 2000, we’ve been striving to make the strongest, most effective developer tools on earth. Today, AI-powered assistance and agents are becoming a core part of how developers work in our IDEs.

We’re building multi-step coding agents that can understand large codebases, plan changes, call tools, and iterate with the user. As a Research Engineer in the Agentic Models team, you’ll be responsible for the models, training loops, and evaluation pipelines that power these agents.

You’ll work at the intersection of SFT and RL-style post-training, and product-driven evaluation, using our distributed GPU and MapReduce clusters to ship models into JetBrains products.

As part of our team, you will:

  • Design, implement, and maintain SFT and RL post-training pipelines for multi-step coding agents.
  • Train and adapt LLMs for agent workflows, including planning, tool use, and multi-step interactions inside JetBrains IDEs.
  • Build and develop evaluation and simulation environments where coding agents can act, be measured, and compared on realistic developer tasks.
  • Design evaluation frameworks and metrics for agent behavior, analyze traces and logs, and close the loop from evaluation back into training, data, and reward design.
  • Analyze training and evaluation results to propose and implement improvements to model architectures, training recipes, and datasets.
  • Work with large-scale infrastructure, including distributed training on GPU clusters and large MapReduce-style data processing for pre-training and fine-tuning datasets.
  • Collaborate closely with research, product, and infrastructure teams to turn high-level product visions into concrete models, experiments, and shipped features. 

We’ll be happy to bring you on board if you have:

  • Extensive hands-on experience training LLMs (pre-training, fine-tuning, or post-training) in a research or production setting.
  • Deep expertise in modern deep learning frameworks such as PyTorch, and specialized LLM training stacks (e.g. Megatron, NeMo, verl, or similar).
  • Strong theoretical and practical understanding of LLM fundamentals: architectures, tokenization, data pipelines, batching, mixed precision, distributed training, and debugging unstable runs.
  • The ability to own projects end to end, starting from a high-level problem or product pain point and overseeing it through the design, experimentation, implementation, and iteration phases.
  • A product-aware mindset – you care about how developers actually use agents and can translate product needs and failure modes into modeling and evaluation work.
  • At least 3 years of Python experience writing clean, maintainable code in modern ML codebases.

Our ideal candidate would have experience with:

  • ML orchestrators and workflow tools such as Kubeflow, Dagster, Airflow, ZenML, and/or job schedulers like Kubernetes or SLURM.
  • Large-scale data and training pipelines, e.g. MapReduce-style clusters, multi-node GPU training, or workloads on the order of 1M+ CPU/GPU hours.
  • Designing and maintaining evaluation pipelines for LLMs or agents, including metrics, dashboards, experiment tracking, and automated regression checks.
  • AI agent development, such as tool-using agents, planners, or multi-step coding workflows, and familiarity with agentic frameworks or patterns.
  • Experiment tracking and observability using tools like Weights & Biases, MLflow, Langfuse, or similar.
  • Inference optimization and serving optimized models in production.

#LI-KP1

We are an equal opportunity employer

We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.

We process the data provided in your job application in accordance with the Recruitment Privacy Policy.

Сайты компаний
J

Founding ML Engineer (Spectrum)

jetbrainsНа сервисе с: 03.10.26 16:18↑ Вакансия с автоподнятием
Зарплата не указанаКипрВеликобританияГерманияНидерландыИспанияСербияЧехияAmsterdam

Software engineers and AI agents alike suffer from the same problem: finding that one person or place that will answer their tough, specific question. Many solutions promise to solve this with similarity search in vector databases. Unfortunately, finding the answer is often a puzzle with pieces to be collected across a myriad of contradictory sources and cannot be solved without surgical search and careful reasoning. 

Spectrum collects data from an organization’s code, docs, and issues, and organizes knowledge in a unified ontology that AI agents can efficiently search through and reason over. We aim to revolutionize the semantic layer space for software-building organizations and move beyond specs that fall out of sync with code, introducing a living spec – one that’s extracted from the whole system and used to keep it aligned. Spectrum is meant to be the single source of truth for all product and architectural knowledge.

Spectrum is a resident of JetBrains' startup incubator, with startup speed and autonomy, and backed by 25 years of developer tooling expertise. We are looking for a top-class ML Engineer who will help us shape the future of software development. You will own our AI and ML engineering stack and help define the research agenda for our team. Your technical vision and design decisions will directly shape the product and determine its success.

Your responsibilities will include:

  • Designing and building the ML/LLM solution for data ingestion, knowledge extraction, retrieval, and subsequent reasoning.
  • Creating the datasets, metrics, and pipelines that drive measurable improvements across the system.
  • Architecting and improving agents for context retrieval, knowledge extraction, and data alignment, which includes prompt engineering, model selection, and inference optimization.
  • Establishing MLOps practices, including orchestration, observability, and experiment tracking.
  • Collaborating with the engineering team on system design and with JetBrains Research on the research agenda.
  • Defining hiring criteria, growing the ML team, and shaping the ML team culture.

We expect you to have:

  • A proven track record as an ML/AI Lead.
  • At least five years of experience in ML/AI systems, with at least two years focused on LLMs and generative AI.
  • A deep understanding of the LLM ecosystem, including model architectures and fine-tuning approaches.
  • Hands-on experience with:
    • Prompt engineering and LLM pipeline design, including evaluation.
    • Agentic frameworks such as LangChain, LlamaIndex, LangSmith, smolagents, or an equivalent.
    • Vector databases and retrieval-augmented generation (RAG) patterns.
    • Deploying and scaling LLM-powered applications using APIs (e.g. OpenAI or Anthropic) or open-source models.
  • Strong Python skills – Kotlin knowledge would be a plus.
  • Excellent communication skills, with the ability to explain complex technical concepts to diverse audiences.
  • Proficiency in English, both written and verbal. 

Our ideal candidate would have:

  • Experience with ontologies, knowledge graphs, or graph-based reasoning.
  • Experience in early-stage startups – you enjoy the zero-to-one phase.
  • The ability to think strategically about product-led AI, beyond just training models in isolation.
  • A background in code analysis, developer tools, or software engineering research.
  • Experience with multi-agent systems or complex agentic workflows.
  • Actively contributed to relevant open-source projects or publications.

What we offer

  • A competitive salary and JetBrains benefits.
  • A generous runway and corporate resources with startup autonomy.

#LI-KP1

We are an equal opportunity employer

We know great ideas can come from anyone, anywhere. That’s why we do our best to create an open and inclusive workplace – one that welcomes everyone regardless of their background, identity, religion, age, accessibility needs, or orientation.

We process the data provided in your job application in accordance with the Recruitment Privacy Policy.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.