Remote Jobs RockRemote Jobs Rock

Senior Data Engineer

🕒 17 days ago
Data EngineeringPythonSQLApache Spark

📜 Description

  • Design and build scalable, cloud-native data platforms from greenfield to production.
  • Implement near-real-time ingestion pipelines using event-driven patterns.
  • Define and enforce platform standards, including Data Lake / Lakehouse principles, medallion architecture, and data contracts.
  • Refactor and optimise existing Spark and PySpark scripts for performance and maintainability.
  • Introduce best practices for code quality, testing, and CI/CD across data pipelines.

🛠️ Requirements

  • 5+ years of professional experience in Data Engineering
  • Strong Python and SQL development skills for pipeline development and optimisation
  • Proficiency in Apache Spark / PySpark, including query optimisation and performance tuning
  • Hands-on experience with Databricks (preferred) or Snowflake
  • Experience with at least one major cloud provider: Azure (preferred), AWS, or GCP
  • Experience with stream processing technologies (Kafka, Spark Structured Streaming)
  • Solid understanding of ETL/ELT patterns, data modelling (dimensional, Data Vault), and data warehousing
  • Experience with orchestration tools (Apache Airflow, Azure Data Factory, or equivalent)
  • Knowledge of Infrastructure as Code (Terraform or equivalent)
  • Understanding of production-grade system requirements: reliability, scalability, observability, and performance
Full job description

Company Description

Are you passionate about building cutting-edge, AI-ready data platforms from the ground up? We are looking for a Senior Data Engineer to join our Data Engineering Team and lead high-impact, greenfield initiatives.

You will work on building modern cloud-native data platforms, migrating on-premises legacy systems to the cloud, and laying the architectural foundation for AI-ready data infrastructure. 

In this role, you will collaborate closely with Machine Learning, Data Science, and Product teams, serving as a key technical contributor and thought leader. You will also drive R&D efforts around agentic AI architectures, event-driven systems, and LLM-ready data pipelines – turning architectural concepts into production-grade solutions.

Job Description

  • Design and build scalable, cloud-native data platforms from greenfield to production
  • Implement near-real-time ingestion pipelines using event-driven patterns
  • Define and enforce platform standards, including Data Lake / Lakehouse principles, medallion architecture, and data contracts
  • Refactor and optimise existing Spark and PySpark scripts for performance and maintainability
  • Introduce best practices for code quality, testing, and CI/CD across data pipelines
  • Drive adoption of AI tooling and agentic workflows within the data engineering team
  • Ensure data quality, observability, and reliability across all pipelines and platforms
  • Develop self-service tooling and microservices to simplify platform usage for other teams

Qualifications

  • 5+ years of professional experience in Data Engineering
  • Strong Python and SQL development skills for pipeline development and optimisation
  • Proficiency in Apache Spark / PySpark, including query optimisation and performance tuning
  • Hands-on experience with Databricks (preferred) or Snowflake
  • Experience with at least one major cloud provider: Azure (preferred), AWS, or GCP
  • Experience with stream processing technologies (Kafka, Spark Structured Streaming)
  • Solid understanding of ETL/ELT patterns, data modelling (dimensional, Data Vault), and data warehousing
  • Experience with orchestration tools (Apache Airflow, Azure Data Factory, or equivalent)
  • Knowledge of Infrastructure as Code (Terraform or equivalent)
  • Understanding of production-grade system requirements: reliability, scalability, observability, and performance
  • Upper-Intermediate English level

WILL BE A PLUS

  • Familiarity with RAG pipeline design and LLM integration patterns
  • Knowledge of data governance frameworks and tools (Unity Catalog, Apache Atlas, or similar)
  • Experience with dbt for data transformation and modelling
  • Familiarity with MLflow, Feature Stores, or ML platform integration

Additional Information

PERSONAL PROFILE

  • Self-driven and proactive in identifying improvements
  • Comfortable working in a fast-paced, innovative environment
  • Strong problem-solving mindset with attention to detail
  • Open to experimenting with emerging technologies and approaches
Reddit

Staff Data Scientist - Ads Measurement, Signals, Privacy

Reddit👥 1001 - 5000 employees🏢 Online Community/social Media
🔥 10 hours ago

As a Staff Data Scientist, you will take ownership of critical advertising challenges, focusing on measurement, identity resolution, and signal quality to enhance the advertiser experience on Reddit.

Statistical ModelingIdentity ResolutionExperimental DesignCausal Inference
Reddit

Sr. Staff Data Scientist - Ads Measurement, Signals, Privacy

Reddit👥 1001 - 5000 employees🏢 Online Community/social Media
🔥 10 hours ago

Lead the scientific strategy for Reddit's ads measurement and signal systems, focusing on innovation in privacy-aware measurement and causal inference methodologies.

Data ScienceCausal InferenceStatistical ModelingExperimental Design

Trusted by Remote Workers