Remote Jobs RockRemote Jobs Rock

Senior Site Reliability Engineer - Security and Data Systems (Federal)

๐Ÿ“… Feb 12
Site Reliability EngineeringInfrastructure As CodeTerraformCI/CD

๐Ÿ“œ Description

  • Design, build, and maintain core infrastructure for security SaaS offerings, ensuring high availability and performance.
  • Develop robust automation to eliminate manual processes and ensure consistency across environments.
  • Collaborate with security teams to embed a security-first mindset in all processes and infrastructure.
  • Participate in on-call rotations, leading incident response and root cause analysis.
  • Partner with development and data science teams to guide architectural decisions and implement new services.

๐Ÿ› ๏ธ Requirements

  • U.S. Person Status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee)
  • Strong coding skills with experience in writing production-level code.
  • Deep experience with Terraform for infrastructure provisioning and management.
  • Familiarity with CI/CD practices and tools, particularly Spinnaker.
  • Expertise in container technologies and managing Kubernetes clusters.
  • Experience with database schema management tools like Flyway.
  • Direct experience with large-scale data systems, specifically Snowflake.
  • Interest in AI/ML technologies for improving reliability and operational efficiency.
Full job description

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software engineering and systems administration, where you'll be responsible for building and maintaining highly reliable, scalable, and secure infrastructure. You will be a key contributor, applying your expertise to automate manual processes and proactively solve complex problems before they become incidents, handling incidents, and includes on-call shifts.

*This position requires the ability to access U.S. National Security information. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

Responsibilities

  • Platform & Reliability: Design, build, and maintain the core infrastructure that underpins our security SaaS offerings, ensuring high availability, performance, and scalability. This includes building and operating the tooling for our Snowflake data systems.
  • Automation: Develop robust automation using code to eliminate toil and ensure consistency across our environments. You'll be a key driver in automating everything from infrastructure provisioning to application deployment and incident response.
  • Security & Compliance: Work closely with our security teams to embed a security-first mindset into all our processes and infrastructure. You will be responsible for ensuring our systems and data platforms are compliant with industry standards.
  • Incident Response: Participate in on-call rotations and be a primary responder for critical incidents, leading root cause analysis and implementing preventative measures to ensure issues don't recur.
  • Collaboration: Partner with development, data science, and security teams to provide expert guidance on architectural decisions, best practices, and the implementation of new services.

Key Skills & Qualifications

  • U.S. Person Status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee)
  • Strong Coding Skills: You are a developer at heart and are comfortable writing production-level code to solve complex operational challenges.
  • Infrastructure as Code (IaC): Deep experience with Terraform for provisioning and managing cloud infrastructure and services.
  • Continuous Delivery: Familiarity with modern CI/CD practices and tools, particularly Spinnaker, to automate and standardize our release pipelines.
  • Containerization & Orchestration: Expertise in container technologies and hands-on experience managing large-scale, production-ready clusters with Kubernetes.
  • Database Migrations: Experience with database schema management tools like Flyway for safely and reliably handling database changes.
  • Data Systems: Direct experience with large-scale data systems, specifically with the Snowflake platform.
  • AI/ML Experience (a plus): Experience or a strong interest in AI/ML, particularly how these technologies can be applied to improve reliability, security, and operational efficiency (e.g., AIOps, predictive analysis).
  • Problem-Solving: Excellent analytical and problem-solving skills with a proactive approach to identifying and addressing potential issues.

This role requires in-person onboarding and travel to our San Francisco Office during the first week of employment.

#LI-Hybrid
#LI-MA

(P18058_3355591)

Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.okta.com/us.

The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between:
$147,000-$202,400 USD

The Okta Experience

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation.

Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.

Crunchyroll

Senior Site Reliability Engineer

Crunchyroll๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Entertainment
๐Ÿ•’ yesterday

As a Senior Site Reliability Engineer, you will enhance the reliability, scalability, and security of Crunchyroll's data platforms while driving modern SRE practices and cross-functional collaboration.

Site Reliability EngineeringKubernetesGoogle Cloud PlatformInfrastructure As Code
Braze

Senior Site Reliability Engineer I

Braze๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development๐Ÿค B2B
๐Ÿ•’ 4 days ago

As a Senior Site Reliability Engineer, you will ensure the reliability and uptime of internal services, collaborating with engineering teams to enhance infrastructure and automation.

Site Reliability EngineeringInfrastructure As CodeChefTerraform
Godaddy

Staff Site Reliability Engineer-Observability

Godaddy๐Ÿ‘ฅ 5001 - 10,000 employees๐Ÿข Internet
๐Ÿ•’ 4 days ago

As a Staff Site Reliability Engineer specializing in Observability, you will enhance monitoring systems, manage cloud migrations, and ensure compliance across a diverse infrastructure.

AWSKubernetesLinux AdministrationGo
Okta

Staff Site Reliability Engineer - Kubernetes

Okta๐Ÿ‘ฅ 10,000+ employees๐Ÿข Software Development
๐Ÿ•’ 12 days ago

The Site Reliability Engineer will architect and manage Kubernetes platforms on AWS, ensuring high availability, performance, and cost optimization while automating infrastructure management.

KubernetesAWSHelmTerraform

Trusted by Remote Workers