Live opening · Posted 18 hours ago

Principal Software Engineer, Distributed Systems

Sigma Automate · United States (Remote)
Linkedin Yes
You are 18 hours behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 18 hours ago
CompanySigma Automate
LocationUnited States (Remote)
Work modeYes
SkillsPython, Kubernetes, PostgreSQL
SourceLinkedin
Listed18 hours ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
17 min from Linkedin publishing this role to us finding it
15 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,583 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Background
We are hiring a Staff / Principal Site Reliability Engineer focused on distributed systems and platform reliability. This is not a traditional DevOps role.
You will own the architecture and engineering practices that ensure Sigma workflows execute reliably even when individual components fail.
Processes crash. Networks become unavailable. APIs time out. Workers hang. Messages may be delivered more than once. Databases experience contention. Kubernetes pods restart.
Your job is to design Sigma so those conditions are expected, detected, and automatically recovered from.
The fundamental principle is simple:
An accepted Sigma job may fail, but it should never disappear.
You will work closely with engineering leadership to evolve Sigma's execution architecture, improve production reliability, establish observability standards, and eliminate classes of distributed-system failures before they affect customers.
What You'll Own
Design and improve the systems responsible for executing long-running and scheduled infrastructure workflows.
Build patterns for:
Durable workflow execution
Idempotent operations
Automatic retries and exponential backoff
Worker leases and heartbeats
Stuck-job detection and recovery
Workflow timeouts and cancellation
Failure isolation
Dead-letter handling
Checkpointing and workflow resumption
Concurrency control
Distributed locking and lease management
Reconciliation loops
Graceful degradation during dependency failures
Evaluate and implement workflow technologies such as Temporal or equivalent systems where appropriate.
Messaging and Event Infrastructure
Own and improve Sigma's asynchronous execution architecture.
Work extensively with technologies such as:
Kafka
Distributed consumers
Consumer groups
Message delivery semantics
Partitioning
Consumer lag
Backpressure
Retry strategies
Event ordering
Duplicate message handling
Poison-message isolation
Ensure event-processing failures cannot silently cause workflows to stop progressing.
Database Reliability
Help design application patterns that remain reliable under high concurrency.
Deeply understand and troubleshoot:
PostgreSQL transactions
Row and advisory locking
Deadlocks
Lock contention
Transaction isolation
Connection pooling
Long-running transactions
Query performance
Database failover
Schema and migration safety
Partner with application engineers to identify and eliminate patterns that create production contention or reliability risks.
Kubernetes and Platform Reliability
Improve the reliability of Sigma deployments running in containerized environments.
Responsibilities include:
Kubernetes workload architecture
Pod lifecycle and failure recovery
Resource limits and capacity planning
Autoscaling
Health checks
Graceful shutdown
Rolling deployments
High availability
Infrastructure-as-code
Disaster recovery
Backup and restore validation
Design systems that tolerate node, pod, process, and dependency failures without losing customer work.
Observability
Make the state of Sigma understandable in real time.
Build and improve:
Metrics
Distributed tracing
Structured logging
Dashboards
Alerting
Workflow-level telemetry
Kafka consumer monitoring
Database performance monitoring
Worker health monitoring
Queue depth and backlog monitoring
Technologies may include:
OpenTelemetry
Grafana
Prometheus
Sentry
CloudWatch
Engineers should be able to answer questions such as:
Why is this workflow still running?
Which worker owns this task?
When did it last make progress?
Which dependency is causing the delay?
Did this operation execute once or multiple times?
Can the workflow safely resume?
Are scheduled jobs starting when expected?
Reliability Engineering
Help establish engineering practices expected of mature enterprise platforms.
This includes:
Service-level objectives and indicators
Error budgets
Capacity planning
Load testing
Failure-mode testing
Chaos testing
Incident response
Root-cause analysis
Production readiness reviews
Runbooks
Automated recovery
Disaster-recovery testing
Move Sigma from detecting incidents to preventing and automatically recovering from them.
What We're Looking For
You have significant experience operating production distributed systems where reliability matters.
Strong candidates will have deep experience with several of the following:
Kubernetes
Kafka or similar event-streaming systems
PostgreSQL
Distributed systems
Asynchronous worker architectures
Workflow orchestration
Temporal, Cadence, Step Functions, Conductor, or similar systems
Python backend systems
Cloud infrastructure
Infrastructure-as-code
Observability platforms
High-availability architectures
You understand concepts such as:
At-least-once delivery
Idempotency
Leases and fencing tokens
Distributed locks
Leader election
Retry semantics
Backpressure
Eventual consistency
Transaction boundaries
Failure domains
Reconciliation
Split-brain scenarios
Distributed tracing
Most importantly, you have personally diagnosed and improved production systems experiencing real distributed-system failures.
What Success Looks Like
Within your first several months, you will help Sigma establish an architecture where:
Scheduled workflows reliably begin when expected
Accepted work cannot silently disappear
Worker crashes do not result in lost workflows
Infrastructure failures automatically recover where possible
Long-running workflows can safely resume
Duplicate execution is prevented or safely handled
Stuck workflows are automatically detected
Database and messaging contention is visible before becoming an outage
Engineers can trace a workflow from request through every execution step
Customer environments can scale without requiring manual infrastructure babysitting

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App