Live opening · Posted 7 days ago

Principal Cloud Platform Architect - AWS

Logic Hire Solutions LTD · United States (Remote)
Linkedin Yes
You are 7 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 7 days ago
CompanyLogic Hire Solutions LTD
LocationUnited States (Remote)
Salary$150K/yr - $200K/yr
Work modeYes
SourceLinkedin
Listed7 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
24 min from Linkedin publishing this role to us finding it
16 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
71,560 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

About The Role
Loghic Hire is seeking a senior, hands-on architect to lead the containerization of a large-scale API estate for an enterprise financial services client and turn an approved AWS EKS reference architecture into a secure, repeatable production platform.
This is an architect-who-builds role: you will write and review Terraform, Helm, Kustomize, Flux, and Istio policy, and stay close enough to the code, pipelines, and operational data to prove the platform works. Application images stay stateless and portable; the platform owns transport security, identity, routing, secrets, telemetry, admission policy, and rate limiting.
You will operate at the intersection of architecture, hands-on engineering, and technical leadership — defining the target state, then building it, documenting it, and proving it in production. Expect to move fluidly between whiteboarding an ADR, reviewing a Terraform module, debugging an Istio waypoint, and coaching an SRE through a canary rollout.
Key Responsibilities1. EKS Platform Foundation & Golden Path
Design, build, and operate a repeatable, multi-AZ Amazon EKS Auto Mode foundation that serves as the golden path for all application teams.
Implement and harden core cluster add-ons: VPC CNI, Karpenter (node provisioning and scaling), KEDA (event-driven autoscaling), CoreDNS, and metrics-server.
Author reusable Terraform modules, Helm charts, and Kustomize overlays that encode platform standards and can be consumed self-service by application teams.
Establish cluster lifecycle management: version upgrades, node rotation, add-on upgrades, and disaster recovery — with documented runbooks.
Define multi-tenancy boundaries (namespaces, quotas, resource limits, RBAC) and enforce them through policy.
Ensure the platform is fully documented and operable by SRE, including on-call runbooks, dashboards, and escalation paths.
Define and publish golden path templates so new services can be onboarded in hours, not weeks.
GitOps Delivery & Continuous Reconciliation
Make Flux the sole production deployment path — no ClickOps, no manual kubectl applies, no out-of-band changes.
Implement continuous reconciliation across all clusters, with drift detection, alerting, and automated remediation.
Integrate GitLab CI and Artifactory into the build-and-promote pipeline; enforce signed, digest-pinned images end to end.
Build environment promotion workflows (dev → staging → prod) with policy gates and approval controls.
Implement Flagger SLO-gated canary deployments, including metric analysis, automatic rollback, and progressive traffic shifting.
Own the GitOps repository structure, branching model, and secrets-handling strategy (SOPS/Sealed Secrets/External Secrets).
Drive image automation for dependency and base-image updates where appropriate.
Service Mesh — Istio Ambient
Architect and operate Istio Ambient mode: istiod, istio-cni, ztunnel, and opt-in waypoints for L7 policy.
Implement SPIFFE workload identity and strict mTLS across all east-west traffic.
Author and enforce AuthorizationPolicy and default-deny NetworkPolicy for zero-trust segmentation.
Define waypoint placement strategy — when to use L4 ztunnel vs. L7 waypoint — and document the trade-offs.
Measure and report mesh latency and overhead; validate the ambient model meets performance budgets.
Own mesh upgrade strategy, version skew policy, and troubleshooting playbooks.
Provide mesh self-service patterns so application teams can declare their own policies safely.
API Gateway & API Containerization
Deliver a single governed north-south ingress using Tyk Self-Managed, fully Operator-driven from Git.
Implement authentication, authorization, rate limiting, quotas, routing, and API versioning at the gateway.
Enable REST-to-gRPC transcoding so legacy REST consumers can reach modern gRPC services transparently.
Lead the migration of existing APIs onto the platform — inventory, prioritization, cutover planning, and rollback strategy.
Define API onboarding standards and self-service workflows for application teams.
Own gateway observability, analytics (Tyk Pump), and policy lifecycle.
Security, Compliance & Audit
Enforce Pod Security Standards (PSS) across all namespaces, with documented exceptions.
Implement Kyverno admission policy as code — validation, mutation, and generation rules managed via Git.
Configure EKS Pod Identity and least-privilege IAM roles for workloads and add-ons.
Integrate AWS Secrets Manager / CSI driver and KMS for secrets and encryption at rest.
Manage certificate lifecycle with cert-manager (issuance, rotation, revocation).
Implement image signing and SBOM generation, with admission-time verification.
Build a durable billing/audit capture path (Kinesis → Firehose → S3 Object Lock) with an approved, load-tested reliability contract.
Produce threat models and participate in security reviews with the client's risk and compliance teams.
Support PCI-scoped and SOC 2 control objectives across the platform.
API Contracts, Reliability & Operations
Establish gRPC/Protobuf as the east-west standard, with buf breaking-change checks enforced in CI.
Define SLOs and error budgets for platform services and the golden path; report on them regularly.
Instrument the platform with OpenTelemetry for metrics, logs, and traces — built in, not bolted on.
Build dashboards, alerts, and rollout analysis tooling for platform and application teams.
Drive performance validation at tens of thousands of TPS, including load testing and capacity planning.
Own incident response for the platform during build-out, including participation in the escalation rotation.
Continuously improve reliability, cost efficiency, and operational ergonomics.
Technical Leadership & Client Engagement
Author and maintain Architecture Decision Records (ADRs), standards, and reference architectures.
Produce threat models and design reviews for critical platform components.
Coach and mentor platform engineers, SREs, and application engineers — raise the technical bar across teams.
Represent Loghic Hire in client architecture reviews, steering committees, and readiness assessments.
Collaborate with client stakeholders to align platform roadmap with business and regulatory priorities.
Contribute to internal Loghic Hire practice development — reusable patterns, playbooks, and accelerators.
Tech StackCategoryTechnologiesCloud & Container OrchestrationAWS, Amazon EKS (Auto Mode), EC2, VPC, IAM, KMS, Secrets Manager, EKS Pod Identity, Karpenter, KEDAInfrastructure as CodeTerraform, Helm, KustomizeGitOps & CI/CDFlux, GitLab CI, Artifactory, Flagger, image automationService MeshIstio (Ambient mode), istiod, istio-cni, ztunnel, waypoints, SPIFFE, mTLS, AuthorizationPolicyAPI GatewayTyk Self-Managed, Tyk Operator, Tyk Pump, REST-to-gRPC transcodingNetworking & PolicyVPC CNI, NetworkPolicy, Kyverno, Pod Security StandardsSecurity & Compliancecert-manager, image signing, SBOMs, S3 Object Lock, PCI-scoped environments, SOC 2Event & Audit CaptureAmazon Kinesis, Firehose, S3 Object LockAPI ContractsgRPC, Protobuf, buf, HTTP/2Runtime & LanguagesJava 21, Spring Boot, Go (nice to have)ObservabilityOpenTelemetry, Prometheus/AMP, Grafana, Fluent Bit, Splunk, Honeycomb, SLOs/error budgetsCost & OptimizationGraviton, Kubecost, OpenCostRequirements
10+ years in software, infrastructure, or platform engineering, including 5+ years of hands-on production Kubernetes on AWS — EKS architecture, operations, networking, upgrades, scaling, and incident troubleshooting.
Proven track record of standing up EKS platforms end to end and containerizing existing API workloads onto them at enterprise scale.
Strong IaC and Kubernetes configuration skills (Terraform, Helm, Kustomize) and reusable platform-module design, with defensible architecture decisions and cross-team leadership in a regulated environment.
Production service mesh ownership — Istio architecture, policy, rollout, performance, troubleshooting; Ambient mode especially relevant.
Deep GitOps experience with continuous reconciliation, drift management, and environment promotion; able to implement the target model in Flux.
Practical Kubernetes and AWS security: PSS, policy as code, NetworkPolicy, IAM, Pod Identity, KMS, certificate management, image signing, SBOMs.
API gateway and regulated audit/event-ingestion design; able to own a self-managed gateway (direct Tyk experience strongly preferred).
Working knowledge of gRPC, HTTP/2, and Protobuf, plus enough Java 21 / Spring Boot familiarity to review the reference runtime pattern.
Observability depth across metrics, logs, traces, SLOs, and rollout analysis with OpenTelemetry; performance validation at tens of thousands of TPS.
Nice to Have
Tyk Operator/Pump and custom Go or gRPC plugins
Flux image automation
Flagger progressive delivery
Kinesis/Firehose/S3 Object Lock compliance-grade event capture
Graviton, Kubecost/OpenCost and cloud-cost optimization
Fluent Bit, AMP, Splunk, Honeycomb, Grafana

Work arrangement
Yes

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App