As a Data Engineer, you will develop data integration solutions and contribute to the architecture of software products for real-time processing in the vehicle-to-cloud ecosystem.
Data Engineer -Databricks
🕒 7 days ago
PythonPysparkSQLDatabricks
📜 Description
- Design, develop, and maintain scalable ETL/ELT pipelines using Databricks.
- Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs.
- Build data pipelines with Databricks Unity Catalog.
- Implement business logic, data transformations, and dimensional data models.
- Create, schedule, monitor, and optimize Databricks Jobs and Workflows.
- Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver, Gold).
🛠️ Requirements
- Hands-on experience with the Databricks Lakehouse Platform.
- Experience integrating with REST APIs for data ingestion and data export.
- Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension
- Understanding of data warehousing concepts and best practices.
- Knowledge of partitioning, file optimization, Spark performance tuning, and query
- Experience with Git and CI/CD best practices
- Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus.
- Databricks certification is an added advantage.
Full job description
Position Overview
We are looking for a Data Engineer with hands-on experience in building scalable data pipelines and data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate should have strong expertise in Python, PySpark, SQL, Databricks, AWS, and REST API integrations for data ingestion, managing large volumes of data, and data export
ShyftLabs is a growing data product company that was founded in early 2020 and works primarily with Fortune 500 companies. We deliver digital solutions built to help accelerate the growth of businesses in various industries, by focusing on creating value through innovation.
Job Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines using Databricks,
PySpark, and SQL.
● Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs.
● Build data pipelines with Databricks Unity Catalog.
● Implement business logic, data transformations, and dimensional data models.
● Create, schedule, monitor, and optimize Databricks Jobs and Workflows.
● Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver,Gold).
● Ensure data quality through validations, error handling, logging, and monitoring.
● Optimize Spark workloads for performance, scalability, and reliability.
● Collaborate with cross-functional teams to deliver production-ready data solutions.
PySpark, and SQL.
● Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs.
● Build data pipelines with Databricks Unity Catalog.
● Implement business logic, data transformations, and dimensional data models.
● Create, schedule, monitor, and optimize Databricks Jobs and Workflows.
● Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver,Gold).
● Ensure data quality through validations, error handling, logging, and monitoring.
● Optimize Spark workloads for performance, scalability, and reliability.
● Collaborate with cross-functional teams to deliver production-ready data solutions.
Basic Qualification:
Strong expertise in Python, PySpark, and Advanced SQL.
● Hands-on experience with the Databricks Lakehouse Platform.
● Good understanding of Unity Catalog, Delta Lake, Databricks Workflows/Jobs,
Clusters, Notebooks, Repos, and Medallion Architecture.
● Experience integrating with REST APIs for data ingestion and data export.
● Strong knowledge of ETL/ELT development, batch processing, incremental loading,
and data transformation.
● Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension
tables, SCD concepts).
● Understanding of data warehousing concepts and best practices.
● Experience working with structured and semi-structured data (CSV, JSON, Parquet,
Delta).
● Knowledge of partitioning, file optimization, Spark performance tuning, and query
optimization.
● Experience with Git and CI/CD best practices
● Hands-on experience with the Databricks Lakehouse Platform.
● Good understanding of Unity Catalog, Delta Lake, Databricks Workflows/Jobs,
Clusters, Notebooks, Repos, and Medallion Architecture.
● Experience integrating with REST APIs for data ingestion and data export.
● Strong knowledge of ETL/ELT development, batch processing, incremental loading,
and data transformation.
● Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension
tables, SCD concepts).
● Understanding of data warehousing concepts and best practices.
● Experience working with structured and semi-structured data (CSV, JSON, Parquet,
Delta).
● Knowledge of partitioning, file optimization, Spark performance tuning, and query
optimization.
● Experience with Git and CI/CD best practices
Preferred Qualifications:
4+ years of experience in Data Engineering with 2+ years of hands-on Databricks
experience.
● Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus.
● Databricks certification is an added advantage.
experience.
● Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus.
● Databricks certification is an added advantage.
Similar jobs
Search more Data Engineer jobs🕒 4 days ago
🕒 yesterday
Join the Data Foundations team to build and enhance asynchronous data services, ensuring reliable data synchronization across Box's systems and evolving the data infrastructure.
Data Engineer
🕒 5 days ago
Raising The Village👥 51 - 200 employees🏢 Non-profit Organization Management
🕒 5 days ago
The Data Engineer will design, develop, and maintain data pipelines and quality systems to enhance RTV's analytics and machine learning capabilities, directly impacting programmatic reach.
🕒 11 days ago
As a Data Engineer on the Data Science team at OpenTable, you will design, build, and maintain robust data solutions that power diner features, partner integrations, and advanced AI initiatives.
Engenheiro de Dados Pleno (foco em Analytics)
🕒 4 days ago
Afya👥 1001 - 5000 employees🏢 Higher Education
🕒 4 days ago
As a Mid-level Data Engineer, you will transform and model data, connecting business needs to reliable data sources ready for analysis within a multidisciplinary team.
