Remote Jobs RockRemote Jobs Rock

Staff / Principal Data Engineer

๐Ÿ“… Jun 12
Data EngineeringData LakesApache SparkApache Flink
Apply Now

๐Ÿ“œ Description

  • Own the design, build, and operation of the data lake and ingestion platform end-to-end.
  • Build low-latency batch and streaming pipelines to ingest and normalize data from various sources.
  • Establish data quality, freshness, completeness, lineage, and observability for the platform.
  • Build data pipelines that support generative AI, including embedding generation and vector storage.
  • Own deployment, CI/CD, and operational reliability of the platform on Kubernetes.

๐Ÿ› ๏ธ Requirements

  • Extensive experience building and operating large-scale data platforms and data lakes, with comfort working at high data volumes.
  • Deep, hands-on expertise with Apache Spark, Apache Flink and modern big-data systems.
  • Proven command of best practices for building and maintaining data pipelines in both batch and streaming modes.
  • Strong production engineering skills across the full delivery lifecycle, including Kubernetes and CI/CD tooling, with the ability to ship end-to-end.
  • A track record of owning data infrastructure end-to-end with limited supervision.
  • Experience with generative AI and embedding models, including embedding pipelines, vector databases and retrieval.
  • A cybersecurity or threat intelligence background, with hands-on exposure to threat types such as phishing, mobile threats and malware.
  • Familiarity with transaction data and transaction fraud signals.

โœจ Benefits

  • Bonus / commission: 15%
Full job description

About the Role

We are building an AI-native data platform that powers fraud detection and response across 360 Fraud Protection. We are hiring a Staff or Principal Data Engineer to own the data platform and data lake at the heart of that work. You work hands-on and own the domain end-to-end, alongside a small group of senior engineers, data scientists and product partners.

  • Owns the unified data platform and data lake that powers detection and response across 360 Fraud Protection.
  • Every detection model and downstream AI capability depends on this data foundation, which makes it one of the highest-leverage engineering roles on the team.
  • Stronger, broader and more reliable fraud signal directly improves detection accuracy, reduces customer losses and protects brand trust.

Key Responsibilities

  • Own the design, build and operation of the data lake and ingestion platform end-to-end, from architecture through production reliability.
  • Build low-latency batch and streaming pipelines that ingest signals from internal and external sources, normalize them to a common schema, enrich them with context and serve model-ready data to the layers above.
  • Make adding a new data source a routine task rather than a project, so our view of risk keeps widening over time.
  • Establish data quality, freshness, completeness, lineage and observability so the platform is trustworthy enough to automate on top of.
  • Build data pipelines that ground generative AI, including unstructured text and threat intelligence processing, embedding generation, vector storage and retrieval.
  • Own deployment, CI/CD and operational reliability of the platform on Kubernetes.
  • Partner with data science, product and architecture to turn the platform into a shared foundation across 360 Fraud Protection.

Required Qualifications

  • Extensive experience building and operating large-scale data platforms and data lakes, with comfort working at high data volumes.
  • Deep, hands-on expertise with Apache Spark, Apache Flink and modern big-data systems.
  • Proven command of best practices for building and maintaining data pipelines in both batch and streaming modes.
  • Strong production engineering skills across the full delivery lifecycle, including Kubernetes and CI/CD tooling, with the ability to ship end-to-end.
  • A track record of owning data infrastructure end-to-end with limited supervision.

Preferred Qualifications

  • Experience with generative AI and embedding models, including embedding pipelines, vector databases and retrieval.
  • A cybersecurity or threat intelligence background, with hands-on exposure to threat types such as phishing, mobile threats and malware.
  • Familiarity with transaction data and transaction fraud signals.

Compensation

  • Base salary range: $180 โ€“ $270
  • Bonus / commission: 15%

Travel

  • Minimal travel expected. This is an on-site role based in New York City, with 3โ€“4 days per week in the office.

AppGate is An Equal Opportunity/Affirmative Action Employer and a federal contractor subject to the Rehabilitation Act of 1973 and the Vietnam Era Veterans Readjustment Assistance Act of 1974 as amended, and their corresponding regulations. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or veteran status, age or any other federally protected class. Further, AppGate is an affirmative action employer committed to taking positive steps to employ, advance in employment and otherwise afford equal employment opportunity to protected veterans and individuals with disabilities. In furtherance of AppGateโ€™s policy regarding affirmative action and equal employment opportunity, AppGate has developed a written affirmative action program. This program is available for review upon request by any applicant or employee during normal business hours by contacting the companyโ€™s EEO Coordinator.

Lead the development and sustainment of secure, scalable data solutions for direct ink-write 3D-printing technologies in support of NNSA Defense Programs applications.

Data Management SystemsPythonData IntegrationCloud Solutions
Shyftlabs

Lead Data Engineer - Databricks

๐Ÿ•’ 12 days ago
Shyftlabs๐Ÿ‘ฅ 51 - 200 employees๐Ÿข Management Consulting

Seeking a Lead Data Engineer with extensive experience in building scalable data pipelines on the Databricks Lakehouse Platform, focusing on ETL/ELT processes and data integration.

PythonPysparkSQLDatabricks
Sigma Software

Principal Data Platform Engineer (Swedish Ad Platform)

๐Ÿ•’ 23 days ago
Sigma Software๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software

Drive architectural decisions and establish engineering standards as a Principal Data Platform Engineer, building a next-generation AI-first data platform for a global AdTech ecosystem.

Apache IcebergSparkFlinkTrino
Sigma Software

Principal Data Platform Engineer (Swedish Ad Platform)

๐Ÿ“… Sep 4
Sigma Software๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software

Drive architectural decisions and establish engineering standards as a Principal Data Platform Engineer for a next-generation AI-first data platform in the AdTech ecosystem.

Apache IcebergSparkFlinkTrino
Sigma Software

Principal Data Platform Engineer (Swedish Ad Platform)

๐Ÿ“… Sep 4
Sigma Software๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Computer Software

Drive architectural decisions and establish engineering standards as a Principal Data Platform Engineer, building a next-generation AI-first data platform for a global AdTech ecosystem.

Apache IcebergSparkFlinkTrino

Trusted by Remote Workers