Live opening · Posted 2 days ago

Linux Architect

Arcesium · Bangalore | Hyderabad
Instahyre 6-10 yrs
You are 2 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 2 days ago
CompanyArcesium
LocationBangalore | Hyderabad
Experience6-10 yrs
SourceInstahyre
Listed2 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
23 min from Instahyre publishing this role to us finding it
3 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,165 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

You'll own the architecture and health of the Linux side of Developer Compute, the fleet of AWS WorkSpaces, WSL, and Linux config that thousands of engineers build on every day. You'll set the standard for how the team ships Ansible changes through peer review, do the deep OS-internals work that keeps the fleet healthy, and act as the final technical authority when a problem survives first-line and self-service troubleshooting. This is a hands-on senior Unix engineering role: you diagnose from strace and journalctl, design the fix at the platform level, and make sure it's codified so the same problem doesn't recur across the fleet.
Responsibilities:
Architect and own the Linux compute platform and the design decisions behind AWS WorkSpaces, WSL, and the broader Linux fleet that thousands of engineers build on daily.
Set the automation standard: define how configuration, patching, and lifecycle management are codified in Ansible and enforced through engineering review, not just author individual changes.
Act as the final technical authority on hard Unix/identity/infrastructure problems the ones that survive first-line and self-service troubleshooting.
Drive the platform's technical roadmap: OS/region expansion, deprecation of legacy images, and R& D into Unix-ecosystem capabilities the fleet doesn't yet have.
Measured by outcomes, not activity. Within 6-12 months:
Resolve fleet-wide health issues at the root: Investigate unhealthy/stuck/disconnecting WorkSpaces reported daily by end users; stale desktop-session state, broken shell config after a distro migration; disk/memory exhaustion; CloudWatch-reported anomalies; and fix the underlying cause via Ansible rather than a one-off reboot.
Own Ansible for the Linux/WorkSpaces fleet: Author and review playbooks and roles covering package management, DNS/nameserver config, credential-refresh scripts, PAM/Kerberos, and systemd-based job scheduling; land changes through a peer-reviewed merge-request workflow with staged rollout and a rollback plan.
Troubleshoot auth and identity plumbing: Debug Kerberos/KCM cache issues, kinit failures, GSSAPI/publickey SSH auth errors, and STS regional-endpoint assumptions across accounts the kind of problems that show up as "VSCode SSH won't connect" but trace back to credential caches or IAM endpoint config.
Keep package management and toolchains working: Own issues across modern and traditional Linux package managers and internal package repositories, SSL/cert trust issues, GPG signature failures, and repository-trust enforcement changes so day-to-day package installs keep working across every supported distro and macOS/WSL.
Support new region and OS rollouts: Help stand up new WorkSpaces regions and new OS bundles (e. g., newer Ubuntu/Rocky Linux releases), directory setup, bundle config, lifecycle automation, and portal integration and lead migration paths off end-of-life images.
Build observability, not just fixes: Extend monitoring for Ansible run failures and workspace health so issues surface as alerts before they become a wave of support requests; contribute to diagnostic tooling that lets users self-service common problems.
Debug container and dev-tool issues on the fleet: Resolve container networking-at-boot failures, credential-in-container breakage, dev container support, and IDE/SSH integration issues that block engineers mid-task.
Set the platform standard and mentor: Define how root-cause fixes get codified into Ansible and documentation rather than repeated as generic advice; raise the technical bar of the wider team and be the reference point for how hard Unix problems get diagnosed and closed out permanently.
Requirements:
Experience: 5-8+ years in Linux systems engineering, with real ownership of production Linux fleets, not just following documented procedures.
AWS WorkSpaces / VDI depth: Hands-on experience operating and troubleshooting AWS WorkSpaces (or equivalent Linux VDI) directory integration, bundle/image lifecycle, region rollouts, connectivity and performance diagnosis (WSP protocol, CloudWatch metrics).
Advanced Unix/Linux fundamentals: Comfortable reasoning about systemd, PAM, Kerberos/KCM credential caches, SSH auth mechanisms (publickey, GSSAPI), filesystem/permission issues, and using strace/journalctl/perf to get from symptom to root cause, not just restarting the service.
Ansible in production: Strong hands-on Ansible playbooks, roles, idempotency, staged rollout and comfort working through a peer-reviewed merge-request workflow with CI as the primary way changes reach the fleet.
Package ecosystem knowledge: Familiarity with Linux package managers across distros (apt/dnf), Homebrew, and modern declarative package/environment managers (e. g., Nix-based tooling) including the cert-trust, GPG-signing, and repo-config issues that come with managing them at fleet scale.
R& D mindset: A track record of independently digging into unfamiliar problems (e. g., AWS/Rocky Linux bug backports, kernel-level networking issues) rather than waiting for someone else to hand you a fix.
Cloud and directory services foundation: Solid AWS fundamentals (IAM, STS, VPC networking) and enough Active Directory/Kerberos literacy to debug identity issues that span the two.
Communication under pressure: Able to work a live escalation with a blocked engineer without guessing, diagnose, explain, and close the loop; comfortable being the person others turn to when first-line steps don't work.
Education: Bachelor's in Computer Science, Engineering, or related field; equivalent hands-on experience considered.
Nice to Have:
Experience with WSL as a first-class dev environment alongside Linux WorkSpaces.
Exposure to Docker/container networking issues on developer machines (not production container orchestration).
Familiarity with fleet-management/osquery-style tooling and observability platforms (e. g., Datadog, Grafana).
Prior experience supporting a developer-facing platform (as opposed to purely back-office IT); you understand why a broken git clone or IDE SSH connection is a P1 for an engineer mid-sprint.

Experience
6-10 yrs

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App