Staff Site Reliability Engineer
Soundhound
Apply to this job Toronto Canada (remote) remote Until 10/7/2026 12+ years exp First posted August 8, 2026 Last posted August 8, 2026
Job description
Your Career, our Future—Together.
Ready to join something big? At SoundHound AI, we bring voice, generative, and conversational AI together to transform how people interact with products and services. From voice-enabled vehicles to food ordering and customer support, our multilingual, omnichannel technology already impacts hundreds of millions worldwide.
The Opportunity
We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.
What You'll Do
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
- Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
- Partner with engineering teams to optimize performance, cost, and reliability of backend services.
- Drive incident response, post-mortem analysis, and long-term remediation efforts.
- Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
- Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
- Lead department wide compliance (PCI, SOC) initiatives.
What You'll Bring
- 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
- Experience designing and managing high-throughput, distributed systems.
- Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
- Excellent communication skills and a demonstrated ability to mentor engineers.
Preferred Qualifications
- Experience working in a high-velocity, customer-focused environment.
- Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
- Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
- Experience implementing security and compliance best practices in the cloud.
Workplace & Compensation
This role is available throughout Canada.
Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.
#LI-MQ1 #LI-REMOTE
Let's Start the Conversation
Join SoundHound AI and collaborate with colleagues worldwide who are shaping the future of voice AI. Guided by our values—supportive, open, undaunted, nimble, and determined to win—we strive to build breakthrough AI experiences together.
We provide reasonable accommodations for individuals with disabilities throughout the hiring process and employment. To view our job applicant privacy policy, please visit https://static.soundhound.com/corpus/ta/applicantprivacynotice.html.
Discover more about our philosophy, benefits, and culture at https://www.soundhound.com/careers.
***Please beware of agency recruiters falsely stating that they represent SoundHound AI on job posts. Our job post above will note if we are utilizing a specific agency to assist with the search. Our recruiters use @soundhound.com email addresses exclusively.
About this role
Summary
Design, build, and maintain scalable, reliable infrastructure on Google Cloud Platform.
Job title
Staff Site Reliability Engineer
Experience level
12+ years
Minimum experience
12+ years exp
Industry
software
Location requirements
Remote role, based in Toronto, Canada.
Salary
Not specified
Management role
No
Skills & keywords
Required skills
google cloud platformterraformkubernetesmonitoringprometheus
Preferred skills
functional programmingsecuritycompliancehigh-velocity environment
Specializations
cloud computingkubernetesinfrastructure as codemonitoringdistributed systems
Locations
Structured locations inferred from the posting.
Canada
Remote Country
Related searches