As a Data Engineer, you'll design, build, and maintain data pipelines, ensuring reliable access to data for reporting and analytics while collaborating with various teams.
Data Engineer
📜 Description
- Design and deliver batch and streaming data pipelines from field collection platforms into RTV's Databricks warehouse.
- Implement ELT and ETL workflows in PySpark and Databricks SQL using Medallion Architecture principles.
- Contribute to sprint-based delivery cycles, owning pipeline workstreams from design through deployment.
- Build and maintain Delta Lake table structures with strong schema enforcement and ACID-compliant write patterns.
- Support data observability frameworks, covering data freshness, volume anomalies, and schema drift detection.
- Integrate structured data and ML model predictions into unified warehouse layers for Data Scientists.
🛠️ Requirements
- Bachelor's degree in Computer Science, Software Engineering, Data Engineering, or related field preferred.
- 1-3 years of experience in data engineering or related roles.
- Strong hands-on experience with Python, SQL, and ETL/ELT pipelines.
- Experience with PySpark, Databricks, or similar modern data platforms.
- Familiarity with Delta Lake, Unity Catalog, and data observability tools.
Full job description
Job Title: Data Engineer
Department/Group: Venn
Reporting To: Senior Data Scientist
Years of Experience: 1-3 years
Location: Mbarara preferred. We also welcome applicants who would work from RTV's Kampala office and travel to Mbarara occasionally to meet with the team.
Travel Required: up to 20%
About Raising The Village
Raising The Village's mission is to build and strengthen pathways out of ultra-poverty through an approach guided by data, designed with communities, and scaled by partnerships. We work with last-mile communities to implement solutions that raise incomes, sustain impact, and chart sustainable pathways to possibility within 24 months. To date, we have supported more than two million people in Uganda, Rwanda, Tanzania and the Democratic Republic of Congo, with support from our partners and our team in North America. Find out more about our programs and impact at www.raisingthevillage.org.
The VENN department is the data and technology backbone of our organization, connecting advanced analytics and custom software tools with field implementation to ensure data-informed decision-making at every level.
The Opportunity
RTV's data infrastructure is in an active build phase, with a roadmap that is being deliberately accelerated to keep pace with the organization's expanding program footprint across Uganda, Rwanda, and the Democratic Republic of Congo. The data engineering team is growing, and this hire is part of that growth.
You will join a small, senior-led team at a moment when foundational architecture decisions are still being made, pipelines are being built from the ground up, and the systems you contribute to will directly shape how RTV measures impact, deploys AI tools, and scales programmatic reach to millions of people living in ultra-poverty.
Role Description
The Data Engineer is a core builder on RTV's expanding data infrastructure team, reporting directly to the Senior Data Scientist within the VENN department, and responsible for designing, developing, and maintaining the pipelines, warehouse layers, and data quality systems that power RTV's programmatic analytics, machine learning platforms, and field evaluation tools. Operating at the intersection of data platform engineering, ML infrastructure support, and field data integration, this role works across a fast-moving roadmap that spans batch and streaming ingestion, ELT pipeline development, Delta Lake architecture, observability frameworks, and the integration of structured field data with AI model outputs.
Key Responsibilities
Pipeline Development & Delivery
- Design and deliver batch and streaming data pipelines that ingest data from field collection platforms (SurveyCTO, ArcGIS, custom mobile apps) into RTV's Databricks warehouse, working at pace against an accelerated infrastructure roadmap.
- Implement ELT and ETL workflows in PySpark and Databricks SQL, applying Medallion Architecture (Bronze, Silver, Gold) principles to produce clean, versioned, and consumption-ready data layers.
- Contribute actively to sprint-based delivery cycles, taking ownership of pipeline workstreams end-to-end from design through deployment and monitoring.
Delta Lake & Warehouse Architecture
- Build and maintain Delta Lake table structures with strong schema enforcement, ACID-compliant write patterns, and time travel capabilities to support auditability across program datasets.
- Contribute to data modelling decisions including star schema design, SCD patterns, and denormalization trade-offs.
- Support the evolution of Unity Catalog governance structures including lineage tracking, access controls, and dataset documentation as the warehouse scales across new program domains and geographies.
Data Observability & Quality
- Support the maintenance of data observability frameworks across all pipeline stages, covering data freshness, volume anomalies, schema drift detection, and SLA monitoring.
- Build validation and quality checks at ingestion and transformation layers using tools such as Great Expectations, dbt tests, or Databricks-native monitoring capabilities.
- Contribute to structured logging, alerting, and incident response practices that give the team fast, reliable visibility into pipeline health across all environments.
ML & AI Pipeline Support
- Integrate structured household data, image classification outputs, and ML model predictions into unified warehouse layers for consumption by Data Scientists and the WorkMate AI platform.
- Collaborate with ML Engineers and Data Scientists to build and maintain feature engineering pipelines and training data preparation workflows that support RTV's computer vision and adoption scoring systems.
Collaboration & Documentation
- Work closely with the technical team on roadmap prioritization, architectural decisions, and engineering standards as the team scales.
- Partner with Software Engineers, Data Scientists, field evaluation teams, and program staff to understand data requirements and translate them into reliable, well-documented pipeline solutions.
- Maintain thorough documentation of pipeline architectures, transformation logic, data dictionaries, and runbooks to support team growth and organizational knowledge continuity.
Technical Requirements
Education & Experience: A Bachelor's degree in Computer Science, Software Engineering, Data Engineering, Information Systems, Statistics, or a related quantitative field is preferred. Equivalent practical experience through demonstrable project work, open-source contributions, or bootcamp training is equally welcome. Clear evidence of building and shipping production-grade pipelines, including demonstrable ownership from design through deployment with specific examples of integrating, moving, and transforming data at a meaningful scale.
Technical Skills: Candidates should have strong hands-on experience with Python (including Pandas), SQL, and building or supporting ETL/ELT pipelines. Experience with PySpark, Databricks, or similar modern data platforms is highly valued. Exposure to Delta Lake, Unity Catalog, Medallion Architecture, streaming pipelines, data observability, AWS, and BI tools is considered an asset. The successful candidate will continue to build depth in these areas while working closely with the Senior Data Scientist and broader technical team.
We encourage women, people with disabilities and minority groups to apply for this position. RTV is committed to equal opportunities and diversity of perspective at the workplace.
Disclaimer: Raising the Village DOES NOT charge any kind of FEE(s) at whichever stage of the recruitment process
Similar jobs
Search more Data Engineer jobsAs a Data Engineer on the Data Science team at OpenTable, you will design, build, and maintain robust data solutions that power diner features, partner integrations, and advanced AI initiatives.
People Data Engineer - 100% Remote Job.
The People Data & AI Readiness Engineer is responsible for building and maintaining the data architecture that supports people analytics and AI solutions, ensuring data quality and governance.
As a Data Engineer in the Tax Care team, you will design and maintain data pipelines, migrate platforms, and ensure data integrity for accounting and tax processes.
We are seeking a Data Engineer with at least 3 years of experience in data projects, specializing in Snowflake and ideally DBT, to design and implement data pipelines and analytical models.
