Software Engineer, Observability
Retool
Apply to this job San Francisco on site Until 8/21/2026 3+ years exp First posted May 25, 2026 Last posted May 25, 2026
Job description
ABOUT RETOOL
Nearly every company in the world runs on custom software for critical operations like tracking performance metrics, handling support workflows, building admin dashboards, and countless processes you might never have thought of. But most companies don't have the resources to properly invest in these tools, leading to a lot of old, clunky internal software, or worse, teams still stuck in manual and spreadsheet workflows.
AI has changed who gets to build software. The definition of "developer" now includes analysts, operators, and domain experts creating solutions directly—and the tools they reach for are multiplying by the week. That's both an opportunity and a challenge: as more people build with more AI tools, the risk of shipping ungoverned software into production grows just as fast.
At Retool, we're building the platform that makes all of it safe to ship. Build with any AI tool you want, then deploy into one place that connects to your real business data, enforces enterprise policies automatically, and lets teams create once and reuse everywhere with shared, trusted components. The cost of building software has collapsed. The cost of governing it hasn't—and that's the problem we solve.
Developers and domain experts have already automated over 100 million hours of work on our platform, freeing them to focus on creative problem-solving and strategic work that drives real business value. The people closest to the problem can now build the software to solve it, safely, and within enterprise guardrails.
Let's build the future together.
WHAT YOU'LL DO:
In this role, you will build, integrate, and evangelize observability platforms and solutions for our products and internal systems. You will drive adoption of these solutions and ensure they drive value for the company.
Your core responsibility in this role is to build and deploy observability solutions that make our products highly available, scalable, reliable, observable and delight our customers.
IN THIS ROLE, YOU WILL:
- Help build a great product that improves productivity of engineers across the globe by several orders of magnitude
- Design and build observability solutions for collection, delivery, analysis, and visualization of metrics, logs, and traces
- Work with engineers, designers, product managers and customer support to instrument and implement observability into our products and internal apps
- Build orchestration and automation tooling around off-the-shelf solutions (e.g. Datadog, Grafana), as well as build custom solutions that meet our unique needs
- Be involved in the development of scalable, distributed software systems that support globally distributed customer base
- Coach and mentor other SWE; Provide leadership in iteratively defining & refining development processes as the team grows
THE SKILLSET YOU'LL BRING:
- 3+ years of related professional experience, 2+ years working on a mission critical platform with high-availability requirements
- Experience with containerization (e.g. Docker, Kubernetes), infrastructure as code (e.g. Terraform) and observability (e.g. Datadog, Stackdriver, Wavefront, Grafana) stacks
- A strong understanding of system availability, resiliency, and recoverability
- Strong organizational skills with high attention-to-detail and able to work independently with minimal supervision
- Ability to thrive in a high-energy, high-growth, fast-paced, entrepreneurial environment. Willing to learn new skills and implement new technologies
BONUS POINTS:
- Familiarity with TypeScript and Node.js backend development
- Familiarity with React frontend web development
- Experience with observability platforms and tools like Datadog, Grafana, etc..
Retool offers generous benefits to all employees and hybrid work location. For more information, please visit the benefits and perks section of our careers page!
Retool is currently set up to employ all roles in the US and specific roles in the UK. To find roles that can be employed in the UK, please refer to our careers page and review the indicated locations.
About this role
Summary
Build and deploy observability solutions for products ensuring high availability and reliability
Job title
Software Engineer, Observability
Experience level
3+ years
Minimum experience
3+ years exp
Industry
software
Location requirements
hybrid work in San Francisco, US-based preferred
Salary
Not specified
Management role
No
Skills & keywords
Required skills
containerizationDockerKubernetesTerraformDatadogStackdriverWavefrontGrafana
Preferred skills
TypeScriptNode.jsReact
Specializations
observabilitymetricslogstraces
Locations
Structured locations inferred from the posting.
San Francisco, CA, USA
Hybrid City