neoflex

Data Engineer Spark (Офис/Банк) (Python, Java)

ОДИН ИЗ ЛУЧШИХ РАБОТОДАТЕЛЕЙ РОССИИ

neoflex · На сервисе с: 02.10.26 15:38

Зарплата не указанаРоссияМоскваОфис

ОДИН ИЗ ЛУЧШИХ РАБОТОДАТЕЛЕЙ РОССИИ

Мы – Neoflex. Аккредитованная IT компания. За 20 лет работы мы создали 12+ готовых решений для бизнеса, так же занимаемся заказной разработкой программного обеспечения.

Приветствуем на странице нашей компании и благодарим за интерес к вакансии. Будем рады оказаться полезны друг другу.

Сейчас мы расширяем команду, занимающуюся проектом по миграции данных из 87 источников на единую платформу. Команда входит в Центр компетенций BDS - одно из сильнейших подразделений нашей Компании, экспертов в области работы с большими данными, построения DWH и использующих в работе современный стек BigData: Greenplum, Hadoop, NiFi, Hive, Spark/Scala, PostgreSQL, Informatica, Kafka, Apache Airflow, Apache Griffin.

ПРОЕКТ: Суть проектов связана с построением хранилищ, которые построены на различных базах данных Hadoop, GreenPlum, PosgreSQL. Работаем с витринами и занимаемся разработкой ETL процессов.

СТЕК: Spark, Hadoop, Hive, Hue, Jupyter.

Работа из офиса 5/2 Москва ул. Поклонная

Что нужно делать:

  • заниматься определением источников, выявлением проблем с качеством данных в источниках;

  • разработкой схем данных конечных витрин в Hive по утвержденному проекту;

  • маппингами атрибутов источников на приёмник;

  • разработкой ETL-процессов с помощью внутреннего фреймворка;

  • реконсиляцией данных, написанием базовых тестов на качество данных.

Что мы хотели бы видеть:

  • уверенный уровень sql (в пределах стандарта, без привязки к СУБД);
  • представления о java / python;

Желательно:

  • опыт работы с Hive, Spark;
  • опыт работы с Оozie, Airflow;
  • умение работать с Confluence;
  • есть представление о работе с Jira, Git.

Что ты приобретёшь, присоединившись к нам:

  • достойную оплату труда + компенсационные, стимулирующие и мотивационные выплаты, бонусы за участие в реферальной программе;
  • работа в команде профессионалов готовых делиться экспертизой;
  • официальное трудоустройство по ТК РФ, аккредитация IT, расширенный социальный пакет:

✔️ страховка ДМС (с 3-го месяца работы, стоматология, возможность подключения родственников, теле медицина, полис ВЗР),

✔️ сотрудникам со стажем в Neoflex более 3 месяцев при предоставлении листка нетрудоспособности устанавливается доплата до полного заработка за период болезни,

✔️ обучение детей сотрудников ИТ специальностям,

✔️ компенсация затрат на фитнес и занятия английским языком;

  • обеспечиваем техникой для работы (ноутбук, наушники, мышь);
  • профессиональное развитие - в Учебном Центре (курсы по работе с большими данными, видео лекции, тренажеры, карьерный коучинг, лекции, тренинги, конференции, участие в митапах);
  • возможность пройти проф.сертификацию;
  • прозрачную систему карьерного развития Performance Review;
  • персонального наставника с первого дня работы;
  • насыщенную корпоративную жизнь: яркие корпоративы, праздники для детей сотрудников, корпоративные спортивные мероприятия, мотивационные награждения;
  • комфортную атмосферу в филиалах компании в городах: Москва, Санкт-Петербург, Нижний Новгород, Пенза, Воронеж, Саратов, Самара, Краснодар где есть лаунж и фотозоны, вендинги в кухнях, пространство для медитаций и другие секретные места, о которых знают только наши сотрудники.

Мечты и команды работают вместе. Мы будем рады, если ты станешь частью нашей команды! Откликайся ;)

Похожие вакансии Data Engineer

Сайты компаний
V

Senior Engineer - Hyperscale Analytics

veeamНа сервисе с: 03.10.26 21:30
Зарплата не указанаКанадаSan JoseГибрид

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are seeking exceptional Senior Engineer- Hyperscale Analytics to design and scale next-generation data processing and analytics platforms that power OLTP, OLAP, and large-scale distributed data systems. You will build and optimize pipelines and services that handle billions of records daily, enabling real-time transactions, analytical insights, and AI-driven decisioning.

What You’ll Do

  • Transactional & Analytical Systems: Design and implement highly scalable OLTP systems for real-time workloads and OLAP systems for complex analytical queries on massive datasets
  • Distributed Processing: Build, optimize, and maintain large-scale batch and streaming pipelines using frameworks such as Apache Spark, Flink, Presto/Trino, or Kafka Streams
  • System Performance & Scale: Optimize systems for low-latency queries, high-throughput ingestion, and interactive analytics, ensuring seamless performance as data volumes scale to petabytes
  • Data Infrastructure: Develop and integrate with modern storage and processing systems (e.g., Snowflake, BigQuery, Redshift, Cassandra, HDFS, Delta Lake, Iceberg) to support hybrid analytical/transactional workloads
  • Reliability & Observability: Ensure high availability, reliability, and monitoring across large compute and storage clusters with automated failover and recovery
  • Collaboration: Partner with data scientists, ML engineers, and product teams to build unified, secure, and cost-efficient data platforms

What You’ll Bring

  • 6+ years of professional software engineering experience, with a significant portion focused on data infrastructure, distributed systems, or large-scale analytics platforms
  • Deep experience with distributed data processing frameworks (e.g., Spark, Flink, or similar)
  • Strong background in cloud-scale data architecture (AWS, Azure, or GCP) — data lakes, warehouses, streaming platforms
  • Proficiency in one or more of: Python, Java, Scala, or Go
  • Experience with modern data warehouse/lakehouse technologies (e.g., Snowflake, Databricks, BigQuery, Redshift)
  • Solid understanding of data modeling, ETL/ELT design, and pipeline orchestration (e.g., Airflow, dbt)
  • Track record of designing for scale, reliability, and cost efficiency in production systems

Bonus Skills

  • Experience with HTAP (Hybrid Transactional/Analytical Processing) systems or real-time analytics platforms
  • Familiarity with data lakehouse architectures and formats like Parquet, ORC, Delta, Iceberg, Hudi
  • Knowledge of containerized deployments (Docker, Kubernetes) and cloud-native data architectures (AWS Redshift, GCP BigQuery, Azure Synapse)
  • Background in query engine development or contributing to open-source OLAP/OLTP frameworks

 

What you'll get

  • Unlimited paid time off, 12 paid holidays including 4 global VeeaMe Days for self-care and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents
  • Medical, dental, and vision coverage starting on your first day
  • Mental health support, therapy sessions, and digital wellness tools via our Employee Assistance Program
  • 401(k) retirement plan with company matching contributions
  • Fertility, adoption, and surrogacy support through Maven, plus paid volunteer time
  • AirVet: 24/7 virtual veterinary care at no cost
  • Legal services, identity protection, and supplemental health insurance options
  • Tax-advantaged spending accounts for healthcare, dependent care, and commuting
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Pay Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, skills, and geographic zone. Offers are typically made below the midpoint of the range.

In addition to compensation, Veeam provides a comprehensive benefits package, including health coverage, retirement plans, and unlimited time off.

Compensation Range (TTC / OTE)
$234,840—$436,080 USD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
V

Senior Engineer - Access Entitlements

veeamНа сервисе с: 03.10.26 21:17
Зарплата не указанаКанадаBritish Columbia

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About the Role

We are seeking an exceptional Senior Engineer – Access Entitlements to design and scale the systems that discover, normalize, and reason about who has access to what across our customers' entire data estate — on-prem file shares, Active Directory, and SaaS platforms like SharePoint, Google Workspace, and Microsoft 365. You will build the connectors, pipelines, and graph models that turn millions of raw permission grants into an accurate, queryable picture of access across billions of files and identities, powering least-privilege analysis, access certification, and risk detection for enterprise customers

What You'll Do

  • Entitlement Scanning: Build and extend connectors that enumerate identities (users, groups, service accounts, computers) and resource-level permissions (ACLs, role assignments, sharing links) across on-prem and SaaS data sources, handling per-connector quirks like SID resolution, inheritance breakage, and nested group membership
  • Scale & Concurrency: Design entitlement pipelines that safely parallelize across large tenants — correctly handling shared state, batching, and checkpointing so a scan of hundreds of thousands of principals and files can pause, resume, and recover without data loss or duplication
  • Data Modeling: Own the mapping from raw connector output to normalized entitlement records, and from those records into our identity graph — designing node/edge structures that represent principals, resources, and access grants (including sharing links and group-inherited access) in a way that supports fast traversal at scale
  • Multi-Store Architecture: Work across the full data path — Elasticsearch for search-driven access, Databricks/Delta for large-scale analytical joins, and a Neptune-backed graph for relationship queries — making deliberate tradeoffs about what gets synced where, and keeping those stores consistent as data volume grows
  • Entitlement-Activity Correlation: Build the pipelines that join static entitlements against observed activity logs to answer "who can access this, and who actually does" — the core signal behind over-permission detection and access-risk scoring
  • Reliability & Correctness: Instrument and test for the failure modes unique to entitlement data — partial scans, malformed ACLs, race conditions in concurrent principal writes, and inconsistent state between graph and source-of-truth stores

What You'll Bring

  • 6+ years of professional software engineering experience, with meaningful time spent on identity/access systems, security data pipelines, or large-scale distributed data processing
  • Strong Go, Java, or Python skills, with direct experience writing concurrent/parallel data pipelines (goroutines, worker pools, or equivalent) and reasoning about race conditions and shared state
  • Experience with graph databases or graph query languages (Gremlin, Cypher, or similar) and modeling relationship-heavy data
  • Familiarity with Elasticsearch/OpenSearch and at least one large-scale analytical store (Databricks, Snowflake, BigQuery, Redshift)
  • Understanding of identity and access concepts — RBAC, ACLs, group membership

Bonus Skills

  • Experience with Microsoft Graph API, SharePoint REST APIs, or ActiveDirectory/LDAP-based identity systems
  • Familiarity with AWS Neptune, TinkerPop/Gremlin, or other graph-native databases at production scale
  • Background in DSPM, CIEM, IAM governance, or data security posture tooling
  • Experience designing multi-tenant systems with per-tenant data isolation across search, analytical, and graph store

What you'll get

  • Paid vacation starting at 15 days per year and increasing to 20 or 25 days based on tenure, with 4 additional global VeeaMe Days and 24 paid volunteer hours annually through Veeam Cares
  • Paid parental leave that includes 3 weeks for all parents and 12 weeks for birthing parents
  • Medical, dental, and vision coverage from day one
  • Mental health support, therapy sessions, and digital wellness tools
  • RRSP retirement plan with matching contributions
  • Fertility support, plus 24 paid volunteer hours through Veeam Cares
  • AirVet: 24/7 virtual veterinary care at no cost
  • Opportunities to learn and grow through on-demand libraries (LinkedIn Learning, O’Reilly), mentoring, workshops, and learning events like our annual Global Day of Learning

Compensation Transparency

Veeam is committed to pay transparency and equitable compensation. For this role, the compensation range below reflects the expected total target compensation (TTC), inclusive of base pay and a competitive performance-based bonus. For roles with a commission plan, the compensation range represents On Target Earnings (OTE), which includes base salary plus variable commission. When determining compensation, Veeam takes into consideration factors such as experience, education, and skills. Offers are typically made below the midpoint of the range.

Pay Range
$188,200—$349,400 CAD

Veeam Software is an equal opportunity employer and does not tolerate discrimination in any form on the basis of race, color, religion, gender, age, national origin, citizenship, disability, veteran status or any other classification protected by federal, state or local law. All your information will be kept confidential.

Personal data collected during the recruitment process will be processed in accordance with our Recruiting Privacy Notice, which explains how your information is collected, used, and handled in connection with hiring activities. By applying for this position, you consent to this processing. 

By submitting your application, you confirm that the information provided, including any supporting documents, is complete and accurate to the best of your knowledge. Any misrepresentation, omission, or falsification may result in disqualification from consideration or, if discovered after employment begins, termination of employment.

Сайты компаний
A

Data Analytics Engineer

AvrideНа сервисе с: 03.10.26 15:30↑ Вакансия с автоподнятием
Зарплата не указанаСШАAustinУдалёнка

About the Team

Here at Avride, we're building the future with self-driving vehicles and delivery robots. As you can imagine, this creates a massive amount of data, and we need someone to help us connect the dots. This is a great opportunity to join a growing team and have a real, tangible impact on our technology and our success.

About the Role

We’re looking for a Data Analytics Engineer who is skilled in both data engineering and data analysis. Your main goal will be to take the messy, raw data coming from our vehicles, APIs, and databases, and transform it into clean, reliable datasets that everyone from our engineers to our CEO can actually use. You'll be responsible for building the data pipelines that make this happen, and the dashboards that bring the insights to life.

What You'll Do

  • You'll tame our raw data streams, building the ETL pipelines that pull information from all corners of the company into our data warehouse.
  • You'll work with some truly unique datasets—everything from vehicle sensor telemetry and logistics data to the results of our internal system tests.
  • You'll write the smart, efficient SQL queries needed to shape our data and prepare it for analysis.
  • You'll create and manage the Grafana dashboards that our teams rely on to track performance and spot issues.
  • You'll work closely with other teams to figure out what data they need, and then you'll deliver it.
  • You'll also get to experiment with our internal AI and LLM-based tools to find new ways to analyze data and automate insights.

What You'll Need

  • Strong, practical experience with Python and its data libraries (like Pandas, Polars, etc.).
  • Expert-level SQL. You should be very comfortable with complex joins, window functions, and query optimization.
  • A solid background in building and maintaining ETL pipelines using modern tools.
  • Hands-on experience with a BI tool like Grafana, Tableau, or Looker.
  • Familiarity with workflow orchestrators like Airflow, Dagster, or Prefect.
  • A good high-level understanding of how Large Language Models (LLMs) work and an interest in applying them to data problems.
  • A strong sense of ownership and a passion for making sure the data is right.

Nice to Have

  • You've worked with ClickHouse or other modern analytical databases (Snowflake, BigQuery, Redshift).
  • You have experience with vehicle, sensor, or logistics data.
  • You've worked in the autonomous vehicle or robotics industry before.

#LI-MS1

Candidates are required to be authorized to work in the U.S. The employer is not offering relocation, sponsorship, and remote work options are not available.

Avride is an equal opportunity employer and committed to providing reasonable accommodations to qualified applicants and employees with disabilities to ensure they have equal access to employment opportunities. Avride complies with the Americans with Disabilities Act (ADA), if you need a reasonable accommodation to assist with the application or hiring process, or to perform the essential functions of a job, please email jobs@avride.ai.

HireSeeker собирает вакансии со всех площадок и присылает только релевантные. Бесплатно.