Senior Platform / Site Reliability Engineer

Full-time
On-site
$120,000-160,000/yr
Salt Lake City, UT; Lehi, UT

Job description

Senior Platform / Site Reliability Engineer
Salt Lake City
About Metabody
Your body runs continuously. Your care runs on appointments — fifteen minutes against half a million minutes of signal a year. The data piles up in apps nobody reads between visits, and by the time anyone looks, the window has closed.
Metabody closes the loop. We unify every signal — wearables, labs, records — into one longitudinal graph; our AI co-clinician reads it continuously and surfaces what matters; and a licensed provider in any of 50 states is ready to act. Not just a wearable. Not just telehealth. The whole loop, compounding with every outcome.
We're early, with real patients across all 50 states and a founding team that's built, shipped, and scaled before. This is a founding-team hire: you'll own a function from day one and set the patterns the next ten people inherit. We hold a high bar — for precision, for follow-through, and for treating patients like people in command, not cases to manage.
The Role
You own the infrastructure that patient care runs on — and in healthcare, "runs on" means uptime, security, and auditability are the same job. You'll build the platform the whole team ships on, make PHI safe by default, and turn SOC 2 and HIPAA from a scramble into something the infrastructure enforces on its own.
What you'll do
Own our AWS infrastructure end to end as code: Terraform/IaC, ECS and EKS/Kubernetes, network policies, and least-privilege IAM — reproducible, reviewable, and boring in the best way.
Build for compliance by default: hardened Chainguard/FIPS images, encryption in transit and at rest, audit logging, and the guardrails that make a SOC 2 / HIPAA audit a report you generate, not a fire drill you survive.
Stand up observability that catches problems before patients do — Sentry, Grafana, Datadog — with dashboards and alerting the whole team can read.
Own CI/CD on GitHub / GitHub Actions: fast, safe deploys with security checks in the pipeline, not bolted on after.
Lead incident response — triage, mitigation, blameless postmortems — and drive down the recurrence rate, not just the ticket count.
Use AI tools daily to multiply your output while keeping your judgment in the loop. Inquisitive by default: "wouldn't it be cool if we could…" — and then you build it.
What we're looking for
Deep production AWS experience — proven, not theoretical — with Terraform/IaC, ECS, and EKS/Kubernetes at real scale.
Hands-on with compliance-grade infrastructure: hardened Chainguard/FIPS container images, AWS network policies, secrets management, and least-privilege access.
Experience carrying a regulated environment through SOC 2 (and ideally HIPAA) — you've been the person auditors talk to, and you made the platform do the proving.
Strong observability chops with tools like Sentry, Grafana, and Datadog, and CI/CD ownership on GitHub Actions.
A reliability mindset: SLOs, error budgets, and the instinct to automate away toil rather than heroically absorb it.
Critical AI fluency: you reach for AI tools constantly, and you notice when they're wrong.
Values fit: precision, lead by example, follow-through, high trust and agency, own-it, ruthless empathy.
What we offer
Competitive salary, meaningful early-stage equity, full ownership of a function from day one, and excellent health benefits — traditional insurance plus modern perks.


More information

Minimum education level

Bachelor's

Experience level

Senior (5-7 years)

Job skills

AWS

Terraform

ECS

Kubernetes

SOC 2

HIPAA

Sentry

Grafana

Datadog

CI/CD

Certifications

AWS Certified Solutions Architect

Certified Kubernetes Administrator

SOC 2 Compliance Certification

HIPAA Compliance Training

ITIL Foundation Certification

Company overview

company-logo
Metabody AI