Lead and develop high-performing data engineering and machine learning teams, focusing on scalable data pipelines and production ML systems in a fast-paced SaaS environment.
Manager, Technical Operations
📜 Description
- Own the Infrastructure Service Request intake, triage, prioritization, scheduling, and communication process within JIRA.
- Maintain backlog health, workflow governance, intake quality, prioritization discipline, and clear service-level expectations.
- Coordinate maintenance windows, patching, security updates, operational upgrades, and platform initiatives while minimizing disruption.
- Lead, coach, and develop Cloud Platform Engineers through goal setting, regular 1:1s, performance management, and career planning.
- Serve as an escalation point during major incidents, participate in on-call rotations, lead post-mortems, and implement reliability improvements.
🛠️ Requirements
- Bachelor's degree and at least 5 years of experience in Infrastructure, Platform Engineering, Site Reliability Engineering, or related technical operations roles.
- Demonstrated experience managing technical work intake and backlogs using JIRA or similar workflow platforms.
- Strong understanding of operational governance, technical risk assessment, and business impact analysis.
- Excellent communication skills with experience working with senior technical and business stakeholders.
- Strong experience supporting and troubleshooting Microsoft Azure and Kubernetes.
✨ Benefits
- Eligibility for an annual bonus or commission, depending on role and compensation structure.
- Potential eligibility for overtime pay for applicable roles.
- Compensation determined based on factors including skills, experience, qualifications, geographic location
- Opportunity to directly manage and develop a team of Cloud Platform Engineers.
- Leadership role spanning infrastructure operations, cloud engineering, security, compliance, and automation.
- Opportunity to work with modern cloud, Kubernetes, infrastructure-as-code, observability, and AI-assisted operational technologies.
- Focus on operational reliability, with a target of 99.99% uptime.
- Opportunity to contribute to security, compliance, automation, and continuous improvement initiatives.
Full job description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Manager, Technical Operations based in United States.
As a Manager, Technical Operations, you’ll lead the processes that keep infrastructure work predictable, secure, and aligned with business priorities. You’ll own the intake, triage, prioritization, and scheduling of infrastructure service requests while ensuring strong governance and stakeholder visibility. The role combines technical operations leadership with direct people management, giving you responsibility for developing a team of Cloud Platform Engineers. You’ll work closely with platform engineering, information security, product, and business stakeholders to coordinate maintenance, security updates, and larger infrastructure initiatives. As part of an on-call rotation, you’ll help lead major incident response and drive post-incident improvements focused on reliability and operational resilience. This is an opportunity to modernize infrastructure operations through automation, data-driven processes, and AI-assisted tooling while supporting a 99.99% uptime target.
Accountabilities:
- Own the Infrastructure Service Request intake, triage, prioritization, scheduling, and communication process within JIRA.
- Maintain backlog health, workflow governance, intake quality, prioritization discipline, and clear service-level expectations.
- Coordinate maintenance windows, patching, security updates, operational upgrades, and platform initiatives while minimizing disruption.
- Facilitate infrastructure review sessions that establish priorities, decision-making, execution readiness, and stakeholder alignment.
- Ensure infrastructure changes follow established change-management processes, approvals, governance standards, and audit requirements.
- Lead, coach, and develop Cloud Platform Engineers through goal setting, regular 1:1s, performance management, career planning, and hiring.
- Manage team capacity and staffing requirements while balancing operational demand, workload, and professional development.
- Anticipate and escalate infrastructure risks related to capacity, obsolescence, performance, security exposure, and operational bottlenecks.
- Manage and decompose large infrastructure initiatives, routing appropriately scoped work into platform sprint cycles.
- Serve as an escalation point during major incidents, participate in on-call rotations, lead post-mortems, and implement reliability improvements.
- Partner with Information Security on security hardening, PCI and SOC 2 readiness, remediation activities, evidence collection, and audit coordination.
- Use operational metrics such as cycle time, board aging, change success rates, and throughput to identify bottlenecks and improve workflows.
- Standardize documentation, templates, communication practices, and operational processes to improve efficiency and reduce ambiguity.
- Identify opportunities to reduce manual work through automation, AI-assisted ticket triage, intelligent alerting, anomaly detection, and self-service remediation.
- Provide regular operational updates, risk assessments, and strategic insights to senior platform engineering leadership.
- Bachelor’s degree and at least 5 years of experience in Infrastructure, Platform Engineering, Site Reliability Engineering, or related technical operations roles.
- Demonstrated experience managing technical work intake and backlogs using JIRA or similar workflow platforms.
- Experience coordinating maintenance windows, patch cycles, operational readiness activities, and change-management processes.
- Strong understanding of operational governance, technical risk assessment, and business impact analysis.
- Excellent communication skills with experience working with senior technical and business stakeholders.
- Strong experience supporting and troubleshooting Microsoft Azure, including AKS, VNets, NSGs, Load Balancers, VPN/ExpressRoute, managed identities, Azure Policy, and RBAC.
- Strong Kubernetes expertise covering cluster operations, scaling, upgrades, workload security, and troubleshooting.
- Experience with Cloudflare DNS, CDN, WAF, and Zero Trust configurations.
- Experience with Terraform, Ansible, and automation using PowerShell, Bash, and/or Python.
- Experience supporting Windows Server, IIS, .NET Framework/Core, Ubuntu Server, Nginx, and Python Django environments.
- Experience containerizing .NET and Python/Django applications and operating them in AKS with appropriate health probes, scaling, and observability.
- Experience designing and managing firewall rules, NSGs, ACLs, network segmentation, and distributed firewall policies across cloud and on-premises environments.
- Ability to analyze traffic flows, troubleshoot blocked network paths, resolve configuration issues, and enforce least-privilege access.
- Experience with GitHub Actions or Azure DevOps for automated builds, testing, deployments, and environment promotion.
- Strong monitoring, logging, and performance troubleshooting skills using Azure Monitor and Log Analytics; Prometheus, Grafana, or New Relic experience is a plus.
- Familiarity with PCI, SOC 2, and SOX requirements and experience operating cloud environments within compliance frameworks.
- Experience implementing policy-as-code, RBAC standards, secure network design, and infrastructure governance.
- Experience partnering with Information Security teams on hardening initiatives and audit preparation.
- Azure AZ-104, AZ-305, AZ-400, or ITIL 4 Foundation certification is preferred.
- Base salary range of $130,000–$160,000 USD.
- Eligibility for an annual bonus or commission, depending on role and compensation structure.
- Potential eligibility for overtime pay for applicable roles.
- Compensation determined based on factors including skills, experience, qualifications, geographic location, and other job-related considerations.
- Opportunity to directly manage and develop a team of Cloud Platform Engineers.
- Leadership role spanning infrastructure operations, cloud engineering, security, compliance, and automation.
- Opportunity to work with modern cloud, Kubernetes, infrastructure-as-code, observability, and AI-assisted operational technologies.
- Focus on operational reliability, with a target of 99.99% uptime.
- Opportunity to contribute to security, compliance, automation, and continuous improvement initiatives.
Requirements:
Benefits:
Similar jobs
Search more Manager jobsLead a globally distributed software R&D team to develop SaaS applications, data pipelines, and integrations while ensuring alignment with business priorities and engineering excellence.
Lead the Integration, Infrastructure, Security, and Corporate IT teams, focusing on structuring deployment operations and defining infrastructure strategy to enhance efficiency and reliability.
Sr. Manager, End User Services & IT Automation (Hybrid, Bangalore)
Lead IT End User Services and automation strategy in a Mac-first environment, overseeing service desk operations and driving AI integration for enhanced efficiency.
Senior Manager, Global Coverage/Project Manager
The Senior Manager of Global Coverage is responsible for managing a team of Project Managers, ensuring high customer satisfaction, and leading strategic projects in the market research industry.
