Shape foundational infrastructure for large-scale web data products and lead architectural strategy while mentoring engineers in a remote-first environment.
Staff Engineer (Core & MLOps)
📜 Description
- Architect and evolve control and context planes, advancing service and schema registries, SLO enforcement, and operational feedback loops.
- Own the service chassis and golden path, maintaining multi-language Java and Python client libraries and deployment pipelines.
- Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions.
- Operate and improve core platform infrastructure across Kubernetes, Terraform, and event-streaming platforms.
- Lead architectural strategy through Requests for Discussion (RFDs) covering critical platform initiatives.
- Establish reliability engineering practices, including SLOs, SLIs, and fault isolation.
🛠️ Requirements
- 10+ years of experience building scalable distributed backend systems.
- Advanced Java expertise, including reactive frameworks such as Vert.x or Netty.
- Strong Python proficiency and deep experience with gRPC and Protocol Buffers.
- Hands-on production experience with Kubernetes at scale and Terraform.
- Strong reliability engineering background, including SLO/SLI definition and fault tolerance.
- Exceptional technical writing skills and ability to communicate complex concepts.
✨ Benefits
- Fully remote, remote-first working environment with flexible working hours.
- Freedom and flexibility to work from the location where you are most productive.
- Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
- Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
- Opportunities to attend conferences and connect with colleagues across the globe.
- Collaboration with a diverse, multicultural, and globally distributed engineering community.
- High level of autonomy and organizational trust.
- Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Engineer (Core & MLOps) based in Brazil.
This role offers the opportunity to shape foundational infrastructure powering large-scale web data products and distributed engineering teams. You’ll own the architecture of core control and context planes that enable services and AI-driven workflows to operate reliably and efficiently. Working across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud infrastructure, you’ll tackle complex distributed systems challenges at production scale. You’ll establish engineering standards, reliability practices, and service contracts that influence multiple product squads. The role combines hands-on architecture with technical leadership, mentoring, and cross-functional alignment. In a globally distributed, remote-first environment, you’ll have significant autonomy to solve challenging infrastructure problems and influence long-term platform strategy.
Accountabilities:
- Architect and evolve the control and context planes, advancing service and schema registries, SLO enforcement, health-aware routing, automated canary releases, and operational feedback loops.
- Own the service chassis and golden path, maintaining and improving multi-language Java and Python client libraries, standardized workload specifications, Helm charts, and deployment pipelines.
- Define and govern inter-service contracts, including gRPC and Protocol Buffer definitions, API gateway transcoding, versioning policies, and schema evolution standards.
- Operate and improve the core platform infrastructure across Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and database modernization initiatives.
- Lead architectural strategy through Requests for Discussion (RFDs) covering workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and other critical platform initiatives.
- Establish reliability engineering practices, including SLOs, SLIs, error budgets, fault isolation, and automated weighted canary deployments.
- Participate in shared infrastructure on-call rotations, lead incident post-mortems, and convert operational insights into platform improvements.
- Mentor engineers across multiple squads, review architectural proposals, and establish engineering practices that make reliable software development more consistent and efficient.
- 10+ years of experience building scalable distributed backend systems, with a strong track record of creating internal platforms or core libraries adopted across engineering organizations.
- Advanced Java expertise, including reactive frameworks such as Vert.x or Netty, combined with strong Python proficiency.
- Deep experience with gRPC and Protocol Buffers, including schema evolution and backward compatibility in mission-critical systems.
- Hands-on production experience with Kubernetes at scale, Terraform, and event-streaming platforms such as Kafka.
- Experience designing automated telemetry pipelines, materialized views, feature stores, or other feedback systems that use production data to dynamically improve system behavior.
- Strong reliability engineering background, including SLO/SLI definition, blast-radius analysis, fault tolerance, and rigorous service contracts.
- Exceptional technical writing skills and the ability to communicate complex architectural concepts clearly while driving alignment across teams.
- Strong written and interpersonal communication skills suited to a globally distributed, remote-first environment.
- A curious, continuous-learning mindset with an interest in evaluating new technologies, architectures, and engineering approaches.
- Experience with Temporal, DBOS, or similar durable execution platforms is a plus.
- MLOps experience, including model serving, performance monitoring, or production drift detection, is advantageous.
- Familiarity with zero-trust networking and service meshes such as SPIRE, mTLS, Cilium, Istio, or Envoy is beneficial.
- Experience building developer tooling such as CLIs, SDKs, or project generators is a plus.
- Experience with large-scale web scraping or crawling, or contributions to distributed-systems and data-extraction open-source projects, is advantageous.
- Fully remote, remote-first working environment with flexible working hours.
- Freedom and flexibility to work from the location where you are most productive.
- Opportunity to work on core infrastructure supporting large-scale web data pipelines and distributed systems.
- Exposure to cutting-edge open-source technologies, tools, and evolving AI and web data infrastructure.
- Opportunities to attend conferences and connect with colleagues across the globe.
- Collaboration with a diverse, multicultural, and globally distributed engineering community.
- High level of autonomy and organizational trust.
- Opportunities to influence platform architecture, engineering standards, and technical strategy across multiple teams.
Requirements:
Benefits:
Similar jobs
Search more Software Engineer jobsShape foundational infrastructure for large-scale web data products and lead architectural strategies in a remote-first environment as a Staff Engineer (Core & MLOps).
Shape foundational infrastructure for large-scale web data products and lead architectural strategies in a remote-first environment as a Staff Engineer (Core & MLOps).
Shape foundational infrastructure for large-scale web data products and lead architectural strategy in a remote-first environment as a Staff Engineer (Core & MLOps).
Shape foundational infrastructure for large-scale web data products and lead architectural strategies in a remote-first environment as a Staff Engineer (Core & MLOps).
