As a Data Engineer, you will design, build, and enhance modern cloud-based data platforms that power enterprise analytics, reporting, and operational decision-making.
Engenheiro de Dados Pleno
📜 Description
- Design and build ETL/ELT pipelines using Databricks or open-source technologies.
- Ingest, standardize, and enrich unstructured documents such as PDFs and images.
- Curate and anonymize data in accordance with LGPD requirements.
- Build and operate vector indexing pipelines for RAG applications.
- Integrate data platforms with inference services and APIs.
- Automate data workflows through CI/CD pipelines and data quality testing.
🛠️ Requirements
- Advanced Python and SQL skills.
- Solid knowledge of data modeling and medallion/lakehouse architecture.
- Experience building data pipelines for unstructured data.
- Familiarity with vector databases and embeddings.
- Hands-on experience with Databricks, including Delta Lake and PySpark.
- Experience with Git and CI/CD practices.
- Professional experience in cloud environments like Azure, AWS, or GCP.
- Knowledge of LGPD and handling sensitive data.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engenheiro de Dados Pleno based in Brazil.
This role focuses on building reliable, scalable data foundations for modern AI and analytics solutions. You will design and operate ETL/ELT pipelines using Databricks, Delta Lake, and lakehouse architectures. The position combines structured and unstructured data engineering, including the ingestion and preparation of PDFs, images, and text for RAG and evaluation workflows. You will contribute to data curation, anonymization, vector indexing, embeddings, and versioned datasets. The role also involves connecting data platforms with inference services and APIs while balancing performance, cost, monitoring, and reliability. Automation, data quality, CI/CD, and continuous improvement will be central to your work. You will join a dynamic, technology-focused environment where modern cloud and AI practices are used to solve complex business challenges.
Accountabilities:
The role is responsible for designing and operating data pipelines and platforms that support AI-enabled applications, ensuring data quality, governance, scalability, and operational reliability.
- Design and build ETL/ELT pipelines using Databricks or open-source technologies, ingesting and transforming historical data through a bronze, silver, and gold medallion architecture on Delta Lake.
- Ingest, standardize, and enrich unstructured documents such as PDFs and images with appropriate metadata.
- Curate and anonymize data in accordance with LGPD requirements, producing clean datasets for RAG, few-shot workflows, and testing.
- Build and operate vector indexing pipelines, including chunking, embedding generation, incremental index updates, and version control.
- Prepare versioned datasets and golden sets for evaluation frameworks and accuracy benchmarking.
- Structure persistence for feedback cycles, including positive and negative feedback, justifications, and resolution status, supporting quality and SLA dashboards.
- Integrate data platforms with inference services, APIs, and other systems while considering performance, cost, monitoring, and alerting.
- Automate data workflows through Jobs and Workflows, CI/CD pipelines, and data quality testing.
- Support continuous improvement of data engineering practices, reliability, and operational efficiency.
- Advanced Python and SQL skills.
- Solid knowledge of data modeling and medallion/lakehouse architecture.
- Experience building data pipelines for unstructured data, including PDFs, images, and text, and preparing these datasets for RAG applications.
- Familiarity with vector databases and embeddings, such as Databricks Vector Search, pgvector, or similar technologies.
- Experience with data quality, data testing, and versioning practices.
- Hands-on experience with Databricks, including Delta Lake, Workflows/Jobs, notebooks, PySpark, and Spark SQL.
- Experience with Git and CI/CD practices.
- Professional experience working in cloud environments such as Azure, AWS, or GCP.
- Knowledge of LGPD and the appropriate handling of sensitive data.
- Experience with Unity Catalog in governed environments is desirable.
- Familiarity with MLflow, Databricks Model Serving, or Mosaic AI is a plus.
- Experience with Delta Live Tables, Lakeflow, or Auto Loader is desirable.
- Experience in insurance or financial services environments is an advantage.
- Familiarity with Terraform or Databricks Asset Bundles is a plus.
- Databricks Data Engineer Associate or Professional certification is desirable.
- Proactive approach, problem-solving mindset, adaptability, and willingness to continuously learn new technologies.
- Anywhere Office model, providing geographic flexibility depending on project requirements.
- Discount partnerships and partner offers.
- Talent referral program with rewards.
- Birthday day off.
- TotalPass/Gympass.
- Health and quality-of-life initiatives.
- Dynamic technology-focused environment with opportunities to work with modern data and AI technologies.
- Collaborative spaces and recreational environments at company units, including relaxation areas and shared amenities.
- Opportunity to work with Databricks, lakehouse architectures, vector search, RAG, cloud platforms, and AI-focused data engineering.
- Diverse and collaborative environment focused on continuous evolution and knowledge sharing.
Requirements
The ideal candidate combines strong data engineering fundamentals with practical Databricks and cloud experience, plus an understanding of modern data and AI workflows.
Benefits
Similar jobs
Search more Data Engineer jobsThis mid-level data engineering role focuses on building reliable analytical data solutions, ensuring data quality, and collaborating with cross-functional teams.
This role involves building scalable data and machine learning solutions, optimizing data pipelines, and collaborating with technical and business teams in a cloud-based environment.
Transform large-scale regulatory data into actionable insights by designing and implementing advanced data engineering and machine learning solutions.
Ingénieur Intégration de Données de Santé – HL7 / Interopérabilité (H/F)
As a Health Data Integration Engineer specializing in HL7 interoperability, you will design and deploy data exchange interfaces for a clinical decision support platform across Europe.
