Location: Palo Alto, CA
About Xage
Cyberattacks on critical infrastructure, government, and private enterprises are increasing in scale, speed, and impact. Xage Security is the leader in Zero Trust for high-stakes environments, helping organizations protect what matters most. With Xage, organizations can stop cyber threats, strengthen resilience, simplify security, and move faster with confidence.
We have built tremendous momentum across governments and commercial enterprises around the world, and it’s just the beginning. Recognized by Forbes as one of America’s Best Startup Employers, Xage prioritizes creativity, collaboration, and innovation in pursuit of our mission. We are headquartered in Palo Alto, CA and have global teams across North America and EMEA.
We’re passionate about solving problems that have positive, real-world consequences for the lives of everyday people. We hope you’ll join us in protecting what matters most and helping organizations operate with confidence in an increasingly complex threat landscape.
About the Role
We are seeking a Senior DevOps & SaaS Operations Engineer to design, automate, and scale the cloud infrastructure and operational pipelines powering our hosted Zero Trust AI Gateway platform. As our SaaS solution expands across enterprise customers, ensuring high availability, seamless customer deployments, automated upgrades, and proactive monitoring across both cloud and edge environments is critical.
Today, our platform runs on a VM-based and containerized architecture (Docker, systemd, cloud compute instances). In this role, you will own the current infrastructure and automating deployments, tenant management, and monitoring, while designing the architectural bridge and migration path toward Kubernetes (AKS/EKS/GKE) as our scale demands it.
Key Responsibilities
- Customer Deployment & Onboarding Automation: Build and maintain Infrastructure-as-Code (Terraform, Packer) and automated provisioning workflows to seamlessly deploy, isolate, and configure customer environments across cloud VMs and container runtimes.
- Zero-Downtime Upgrade Orchestration: Architect safe, automated upgrade and rollback pipelines for both hosted cloud control planes and customer-side edge gateways/agents, ensuring continuous operations across version shifts.
- Kubernetes Migration Strategy: Lead the future-state container orchestration roadmap—designing, prototyping, and executing the transition from VM/Docker deployments to Kubernetes (EKS/GKE) without interrupting customer SLAs.
- Monitoring, Observability & Alerting: Design and manage unified observability platforms (e.g., Prometheus, Grafana, Datadog, ClickHouse, OpenTelemetry) to track system health, gateway latency, error budgets, and service SLIs/SLOs.
- Ongoing Maintenance & Site Reliability: Drive operational excellence, capacity planning, backup/disaster recovery, patch management, and incident response automation to maintain enterprise-grade uptime SLAs (99.99%).
- Cloud & Edge Security Infrastructure: Implement security best practices across VM images, container registries, secret management systems, and network perimeter controls across AWS/GCP/Azure environments.
Required Qualifications
- Engineering Degree or equivalent experience
- Production SaaS Experience: 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Operations managing commercial SaaS platforms.
- Infrastructure as Code (IaC) & Cloud: Expert-level mastery of Terraform (or OpenTofu), Packer, and deep experience operating within primary cloud providers (AWS, GCP, or Azure).
- VM & Container Management: Hands-on experience deploying, managing, and hardening VM-based workloads (Azure VM/EC2/GCE) along with containerization (Docker, Docker Compose).
- Kubernetes Expertise: Hands-on experience with Kubernetes (EKS/GKE) and Helm, with a clear understanding of how to architect containerized applications for future Kubernetes migration.
- CI/CD & Release Engineering: Proven track record building robust deployment pipelines (GitHub Actions, Ansible, Packer) featuring blue-green deployments, canary releases, and automated testing gates.
- Observability & Logging: Strong hands-on experience configuring distributed tracing, metrics aggregation, and log ingestion stacks (e.g., Prometheus, Grafana, OpenTelemetry, ClickHouse, ELK/Datadog).
- Scripting & Automation: High proficiency in Python, Bash, or Go for operational tooling, dynamic automation scripts, and API integrations.
Preferred Qualifications
- Edge & Distributed Deployments: Experience managing hybrid deployment models where a central SaaS control plane orchestrates distributed edge gateways or on-premises agents/VMs.
- Security & Compliance: Familiarity with SOC 2 Type II compliance, ISO 27001, secret management tools (such as HashiCorp Vault), and zero-trust networking principles.
- Eventing & Service Coordination: Operational experience with Kafka, Consul, or gRPC-based microservice architectures.
Perks
- Salary Range: $140,000 – $180,000 p/yr + Equity
- Full health, dental, vision insurance
- We will process visa transfers and immigration
- Work with founders and executives closely and participate in all aspects of company building
- Early stage opportunity in a massive sized market with proven traction and growing rapidly
Recognition & Momentum
Xage Security has experienced explosive growth and received numerous awards and recognition, including:
- Named by Forbes one of America’s Best Startup Employers 2024-2026
- $17 million contract awarded by U.S. Space Force’s Space Systems Command (SSC) to offer its zero trust access control
- Named in Gartner research on Cyber-Physical Systems Protection Platforms, Zero Trust Network Access, Privileged Access Management, and CPS Secure Remote Access
- Named in Forrester research on Operation Technology Security, IoT, and Microsegmentation
- Named a Gold winner for Identity & Access Security Solution in the 2024 American Business Awards
- Named a Top 10 Security Solution in the CRN Internet of Things 50 list 2024
- ISO 27001:2022, IEC 62443, and FIPS 140-2 Certified
