Senior SRE / Production Reliability Engineer

Symphony Solutions

Apply to this job
Ukraine remote Until 9/26/2026 First posted July 28, 2026 Last posted July 28, 2026
Job description

We're building a multi-brand iGaming platform that helps operators run their business in regulated markets worldwide. It brings together three main products — Player Account Management (PAM), Sportsbook, and Casino — all built on a modern microservices architecture using Scala, with a React/TypeScript back office.

We're looking for a Senior SRE to join our team and help us build and maintain reliable, observable, and scalable infrastructure. You'll work closely with developers, own reliability practices, and contribute to the team's DevOps culture.

Requirements

Must Have: Technical Skills

  • Linux — confident troubleshooting in terminal, logs, processes, networking basics, resource usage.
  • Kubernetes / GKE — strong hands-on understanding of workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
  • GCP — practical experience with cloud infrastructure, especially around GKE, IAM, networking, load balancing, artifact registry, and production diagnostics.
  • Docker / Containers — images, registries, runtime debugging, container lifecycle.
  • Helm — understands Helm releases, values, deployment state, and rollback.
  • GitOps / FluxCD — able to understand and troubleshoot GitOps deployment flow, drift, image automation, and Git-based rollback.
  • CI/CD understanding — can investigate failed pipelines and deployment issues; does not need to be the main pipeline builder if DevOps owns that.
  • Observability — Prometheus, Alertmanager, Grafana; understands metrics, logs, traces, dashboards, alerting, and production monitoring.
  • Networking fundamentals — DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules.
  • SLI / SLO / SLA — can define and apply reliability targets, not just explain the terms.
  • Incident response — experience with production incidents, rollback, service recovery, RCA/postmortem.
  • Databases / messaging operational basics — PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar; enough to monitor health, migrations, indexes, topics, and failure signals.
  • Security / compliance basics — secrets, IAM/RBAC, audit trail, release/change evidence.

Must Have: Reliability / Operational Skills

  • Can assess if a release is technically safe to proceed.
  • Can define what should be monitored during and after deployment.
  • Can verify production health after release.
  • Can prepare or validate rollback plans.
  • Can lead or support post-release watch.
  • Can improve runbooks and incident response procedures.
  • Can identify gaps in dashboards, alerts, deployment checks, and release evidence.
  • Can work with DevOps without duplicating their ownership.

Must Have: Soft Skills

  • Ownership and accountability — follows production issues through to closure.
  • Calm under pressure — can operate clearly during incidents, failed deployments, and rollbacks.
  • Clear communication — gives concise updates: impact, current action, ETA, risk, decision needed.
  • Structured troubleshooting — breaks issues down by app, infra, network, database, config, release diff, external dependency.
  • Risk awareness — can say when a release is risky and explain why.
  • Collaboration — works well with Dev, QA, Product, Support, Release Manager, and DevOps.
  • Documentation discipline — maintains runbooks, rollback steps, incident timelines, and operational evidence.
  • Blameless mindset — focuses on root cause and prevention.
  • Proactivity — finds missing alerts, dashboards, runbooks, and process gaps before incidents.
  • Prioritization — separates real production impact from alert noise.

Nice To Have

  • Terraform / IaC experience.
  • Advanced GCP infrastructure design.
  • ArgoCD or other GitOps tools.
  • Advanced database or performance tuning.
  • Service mesh experience.
  • On-call rotation experience with PagerDuty/Opsgenie or similar.
  • Experience facilitating RCA/postmortem sessions.
  • Betting/gaming domain experience.
  • Good written English for incident updates, release notes, and audit evidence.
  • Basic understanding of AI-related concepts: agents, skills, MCP.

Responsibilities

  • Own production reliability practices together with DevOps and engineering teams.
  • Monitor production health and improve observability.
  • Support releases, hotfixes, rollback, and post-release watch.
  • Validate deployment readiness and rollback readiness.
  • Participate in incident response and RCA/postmortem.
  • Maintain runbooks, dashboards, alerting rules, and operational documentation.
  • Ensure release evidence is traceable and audit-ready.
  • Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.

About this role

Summary

Maintain reliable, observable, scalable infrastructure; support releases, incident response, and documentation.

Job title

Senior SRE / Production Reliability Engineer

Experience level

senior level

Industry

gaming

Location requirements

Ukraine, remote work not specified

Salary

Not specified

Management role

No

Skills & keywords

Required skills

LinuxKubernetesGCPDockerHelmGitOpsCI/CDPrometheusGrafananetworkingSLI/SLO/SLAincident responsePostgreSQLsecurity

Preferred skills

TerraformArgoCDdatabase tuningservice meshon-call experienceRCA facilitationgaming domainAI basics

Specializations

LinuxKubernetesGCPDockerHelm
Locations

Structured locations inferred from the posting.

Ukraine

Remote Country