агентство цифровой трансформации

Data Engineer (Python)

Осуществлять разработку и поддержку процессов ETL/ELT для загрузки данных в базы данных.

агентство цифровой трансформации · На сервисе с: 28.09.26 14:43

Зарплата не указанаБеларусьМинскОфис

Ключевые задачи:

Осуществлять разработку и поддержку процессов ETL/ELT для загрузки данных в базы данных.
Создавать структуры баз данных и определять параметры хранилищ данных.
Осуществлять автоматизацию процессов обработки и загрузки данных.
Осуществлять анализ данных с очисткой от ошибок и дубликатов.
Осуществлять подготовку датасетов из больших объемов данных для анализа.
Выполнять организацию инфраструктуры для хранения и управления данными.
Осуществлять построение конвейеров обработки данных.
Выполнять настройку среды и обучение моделей машинного обучения на подготовленных данных.

Мы ожидаем:

Знание Python или R, SQL.
Понимание работы с API, процессов ETL/ELT, аутентификации и авторизации.
Знание инструментов и технологий Big Data.
Опыт работы с контейнеризацией (Docker).
Знание систем контроля версий (Git/GitLab).
Опыт работы с Linux системами.
Знание BI-инструментов.
Опыт работы с облачными сервисами.
Опыт работы с брокерами сообщений (Kafka, RabbitMQ).

Будет преимуществом:

Опыт применения инструментов машинного обучения и нейронных сетей.
Практический опыт работы с BI-инструментом Apache SuperSet.
Знание Apache Druid, Arenadata DB, Apache Cassandra.

Мы предлагаем:

Финансовая мотивация и гарантии:
Конкурентная заработная плата
Официальное трудоустройство (ТК РБ), полный соцпакет
28 дней отпуска + матпомощь на оздоровление
Медицинская страховка

Профессиональное развитие:
Оплата курсов, конференций и семинаров

Корпоративная культура:
Современный офис в центре Минска
Качественное оборудование
График работы (пн.–чт. 9:00–18:00, пт. 9:00–16:45)
Корпоративы, подарки сотрудникам на значимые даты
Уютные зоны отдыха с игровыми пространствами
Корпоративная программа занятий спортом (All Sports)

Похожие вакансии Data Engineer

Сайты компаний
V

Director, Graph Databases

veeamНа сервисе с: 03.10.26 21:43
Зарплата не указанаСШАКанадаSan Jose

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

You’ll lead two core systems inside Veeam Data Command Center: the Knowledge Graph and the hyperscale data lake integrations. Together, they help customers understand where sensitive data lives, who can access it, how it moves, and whether AI models trained on it can be trusted. 

You’ll lead multiple teams building a searchable, security-aware graph that works at enterprise scale. This role is for a hands-on technical leader who can set clear direction, grow strong teams, and deliver reliable systems—while building an AI-first engineering culture with high standards for quality and security. 

What You’ll Do

  • Set the technical vision and end-to-end architecture for the Knowledge Graph, including the data model, storage engine, and query layer at very large scale 
  • Guide the evolution of the graph schema for data sources, identities, access, classifications, and lineage (property graph and/or RDF) using Amazon Neptune and/or Neo4j 
  • Own the strategy for hyperscale lake and lakehouse integrations, including connectors and scanning engines that ingest metadata and lineage from Delta Lake, Iceberg, Parquet/Avro, and platforms like Azure Data Lake, AWS S3/Glue, and BigQuery without disrupting production 
  • Drive performance and reliability, including standards for indexing, partitioning, and query planning, and tuning traversals and queries (Gremlin, Cypher/openCypher, SPARQL) 
  • Build and scale an AI-first engineering approach where teams use tools like Claude Code, Cursor, and Copilot responsibly, with guardrails for security, maintainability, and code quality 
  • Invest in reusable engineering building blocks (including “Claude skills” and agent workflows) that make teams faster and more consistent 
  • Own delivery outcomes: roadmap execution, operational readiness, incident learning, and cross-team alignment 
  • Hire, coach, and develop leaders, including engineering managers and senior/staff engineers, with clear expectations and growth paths 

What You’ll Bring

  • 10+ years of software engineering experience in data infrastructure, graph systems, or distributed data platforms 
  • 4+ years of engineering leadership experience, including leading through managers and scaling multiple teams 
  • Strong production experience with Amazon Neptune and/or Neo4j, including scaling, operations, and trade-offs (property graph vs. RDF) 
  • Proven ability to lead graph modeling for complex domains, including lineage and permissions at enterprise scale 
  • Deep knowledge of Gremlin, Cypher/openCypher, and/or SPARQL, including performance tuning and query design best practices 
  • Experience with data lakes/lakehouses (Delta Lake, Iceberg, Parquet) across major cloud platforms (Azure Data Lake, AWS S3/Glue, BigQuery) 
  • Experience designing and operating distributed systems using tools like Spark, Flink, or Presto/Trino, with strong judgement on scalability and cost 
  • Strong backend background in Go and/or Python, with the ability to review designs, guide decisions, and unblock teams 
  • Practical experience using AI-assisted development tools and the ability to set standards that keep AI-assisted code secure and high quality 

Bonus Skills

  • Experience operating graph systems at massive scale 
  • Background in data security, access governance, and policy controls 
  • Experience with AI/ML governance tools and practices (e.g., MLflow, Databricks Mosaic AI) 
  • Experience building custom agents, MCP-based workflows, or reusable engineering automation 
  • Infrastructure-as-Code experience (e.g., Terraform or Pulumi) 
  • Contributions to graph standards or communities (GQL, openCypher, SPARQL) 

 

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Pay Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

Compensation Range (TTC / OTE)
$382,560—$710,400 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
V

Senior Engineer - Hyperscale Analytics

veeamНа сервисе с: 03.10.26 21:30
Зарплата не указанаСШАКанадаSan JoseГибрид

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are seeking exceptional Senior Engineer- Hyperscale Analytics to design and scale next-generation data processing and analytics platforms that power OLTP, OLAP, and large-scale distributed data systems. You will build and optimize pipelines and services that handle billions of records daily, enabling real-time transactions, analytical insights, and AI-driven decisioning.

What You’ll Do

  • Transactional & Analytical Systems: Design and implement highly scalable OLTP systems for real-time workloads and OLAP systems for complex analytical queries on massive datasets
  • Distributed Processing: Build, optimize, and maintain large-scale batch and streaming pipelines using frameworks such as Apache Spark, Flink, Presto/Trino, or Kafka Streams
  • System Performance & Scale: Optimize systems for low-latency queries, high-throughput ingestion, and interactive analytics, ensuring seamless performance as data volumes scale to petabytes
  • Data Infrastructure: Develop and integrate with modern storage and processing systems (e.g., Snowflake, BigQuery, Redshift, Cassandra, HDFS, Delta Lake, Iceberg) to support hybrid analytical/transactional workloads
  • Reliability & Observability: Ensure high availability, reliability, and monitoring across large compute and storage clusters with automated failover and recovery
  • Collaboration: Partner with data scientists, ML engineers, and product teams to build unified, secure, and cost-efficient data platforms

What You’ll Bring

  • 6+ years of professional software engineering experience, with a significant portion focused on data infrastructure, distributed systems, or large-scale analytics platforms
  • Deep experience with distributed data processing frameworks (e.g., Spark, Flink, or similar)
  • Strong background in cloud-scale data architecture (AWS, Azure, or GCP) — data lakes, warehouses, streaming platforms
  • Proficiency in one or more of: Python, Java, Scala, or Go
  • Experience with modern data warehouse/lakehouse technologies (e.g., Snowflake, Databricks, BigQuery, Redshift)
  • Solid understanding of data modeling, ETL/ELT design, and pipeline orchestration (e.g., Airflow, dbt)
  • Track record of designing for scale, reliability, and cost efficiency in production systems

Bonus Skills

  • Experience with HTAP (Hybrid Transactional/Analytical Processing) systems or real-time analytics platforms
  • Familiarity with data lakehouse architectures and formats like Parquet, ORC, Delta, Iceberg, Hudi
  • Knowledge of containerized deployments (Docker, Kubernetes) and cloud-native data architectures (AWS Redshift, GCP BigQuery, Azure Synapse)
  • Background in query engine development or contributing to open-source OLAP/OLTP frameworks

 

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Pay Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

Compensation Range (TTC / OTE)
$234,840—$436,080 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
V

Senior Engineer - Access Entitlements

veeamНа сервисе с: 03.10.26 21:17
Зарплата не указанаКанадаBritish Columbia

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are seeking an exceptional Senior Engineer – Access Entitlements to design and scale the systems that discover, normalize, and reason about who has access to what across our customers' entire data estate — on-prem file shares, Active Directory, and SaaS platforms like SharePoint, Google Workspace, and Microsoft 365. You will build the connectors, pipelines, and graph models that turn millions of raw permission grants into an accurate, queryable picture of access across billions of files and identities, powering least-privilege analysis, access certification, and risk detection for enterprise customers

What You'll Do

  • Entitlement Scanning: Build and extend connectors that enumerate identities (users, groups, service accounts, computers) and resource-level permissions (ACLs, role assignments, sharing links) across on-prem and SaaS data sources, handling per-connector quirks like SID resolution, inheritance breakage, and nested group membership
  • Scale & Concurrency: Design entitlement pipelines that safely parallelize across large tenants — correctly handling shared state, batching, and checkpointing so a scan of hundreds of thousands of principals and files can pause, resume, and recover without data loss or duplication
  • Data Modeling: Own the mapping from raw connector output to normalized entitlement records, and from those records into our identity graph — designing node/edge structures that represent principals, resources, and access grants (including sharing links and group-inherited access) in a way that supports fast traversal at scale
  • Multi-Store Architecture: Work across the full data path — Elasticsearch for search-driven access, Databricks/Delta for large-scale analytical joins, and a Neptune-backed graph for relationship queries — making deliberate tradeoffs about what gets synced where, and keeping those stores consistent as data volume grows
  • Entitlement-Activity Correlation: Build the pipelines that join static entitlements against observed activity logs to answer "who can access this, and who actually does" — the core signal behind over-permission detection and access-risk scoring
  • Reliability & Correctness: Instrument and test for the failure modes unique to entitlement data — partial scans, malformed ACLs, race conditions in concurrent principal writes, and inconsistent state between graph and source-of-truth stores

What You'll Bring

  • 6+ years of professional software engineering experience, with meaningful time spent on identity/access systems, security data pipelines, or large-scale distributed data processing
  • Strong Go, Java, or Python skills, with direct experience writing concurrent/parallel data pipelines (goroutines, worker pools, or equivalent) and reasoning about race conditions and shared state
  • Experience with graph databases or graph query languages (Gremlin, Cypher, or similar) and modeling relationship-heavy data
  • Familiarity with Elasticsearch/OpenSearch and at least one large-scale analytical store (Databricks, Snowflake, BigQuery, Redshift)
  • Understanding of identity and access concepts — RBAC, ACLs, group membership

Bonus Skills

  • Experience with Microsoft Graph API, SharePoint REST APIs, or ActiveDirectory/LDAP-based identity systems
  • Familiarity with AWS Neptune, TinkerPop/Gremlin, or other graph-native databases at production scale
  • Background in DSPM, CIEM, IAM governance, or data security posture tooling
  • Experience designing multi-tenant systems with per-tenant data isolation across search, analytical, and graph store

What you'll get

  • Paid vacation starting at 15 days per year and increasing to 20 or 25 days based on tenure, with 4 additional global VeeaMe Days and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave that includes 3 weeks for all parents and 12 weeks for birthing parents
  • Medical, dental, and vision coverage from day one
  • Mental health support, therapy sessions, and digital wellness tools
  • RRSP retirement plan with matching contributions
  • Fertility support, plus 24 paid volunteer hours through Veeam Cares
  • AirVet: 24/7 virtual veterinary care at no cost
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Compensation Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, and skills. Offers are typically made below the midpoint of the range.

Pay Range
$188,200—$349,400 CAD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.