Live opening · Posted 11 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
The candidate will have responsibilities across the following functions:
Technical Architecture and System Design:
Own system-wide architecture across all five engineering streams- every major design decision, from database schema choices to service boundary definitions to API contracts, is yours to lead and document.
Lead structured system design reviews for all significant features and platform changes- producing Architecture Decision Records (ADRs) with explicit trade-offs, not informal Slack threads.
Design and evolve a microservices architecture that is loosely coupled, independently deployable, and operationally simple- avoiding both monolith sprawl and microservice overengineering.
Architect event-driven systems using Kafka: define topic naming conventions, partitioning strategies, consumer group design, schema evolution via Schema Registry, and CDC patterns with Debezium- and enforce these standards across all teams.
Own the polyglot stack decisions: Node.js and Rust for services, React/Next.js for frontends, MongoDB and ClickHouse for data, Kafka for streaming, Apache Camel for integration routing, and OpenSearch for search- choosing the right tool deliberately, not by default.
Define integration architecture standards for API integrations, BOT-based data acquisition pipelines, and ERP/CRM connectors- ensuring every external data flow is reliable, observable, and recoverable.
Evaluate new technologies with an engineering discipline- prototype, benchmark, document the trade-off, and only then commit the team.
Hands-On Engineering (60%):
Write, review, and refactor production code across the stack- you are a contributor, not just an approver.
Conduct architecture reviews and deep code reviews for high-stakes changes- setting the quality bar through your own output, not just your feedback.
Personally own and resolve the hardest engineering problems: cross-service consistency, high-throughput pipeline bottlenecks, cold-start latency on serverless inference paths, and integration reliability under load.
Pair with engineers on complex implementations- your presence in the code is how you transfer knowledge and raise the team's ceiling.
Stay current across the stack- you can context-switch from a Rust service's memory model to a Next.js hydration bug to a Kafka consumer lag issue in the same day.
DevOps, Reliability and Cloud Operations:
Own reliability, security, performance, and cost across all AWS infrastructure- ECS/EKS, RDS, S3 Lambda, CloudFront, API Gateway, and the services that glue them together.
Establish and enforce DevOps and SRE practices: CI/CD pipelines (GitHub Actions), Infrastructure as Code (Terraform or CDK), container orchestration (Kubernetes on EKS), and on-call rotation.
Define and track SLOs and error budgets for all production services- and hold teams accountable to them, not just measure them.
Build observability into every system from day one: structured logging, distributed tracing (OpenTelemetry), and metrics dashboards (Prometheus/Grafana or CloudWatch)- alerting that fires before customers notice.
Lead incident response: own the runbook culture, drive blameless post-mortems, and ensure every incident produces a concrete engineering improvement.
Drive cost discipline on AWS- right-size instances, optimise data transfer costs, benchmark self-hosted open-source options against managed services, and report on infra spend monthly.
Data Engineering and Ingestion Infrastructure:
Partner with the Data Engineering team to ensure the full data lifecycle is production-grade: raw ingestion from distributor APIs, scraping pipelines, and partner feeds; transformation and enrichment via Spark/PySpark and dbt; storage in MongoDB, ClickHouse, and OpenSearch; and delivery of clean, analytics-ready datasets to AI models and dashboards.
Own the architecture for high-volume, high-frequency data ingestion- the platform ingests component availability, pricing, and supplier data continuously; pipeline reliability and freshness are product features, not just engineering concerns.
Review and guide data model design across MongoDB (operational), ClickHouse (analytics), and OpenSearch (search)- balancing write performance, query latency, and schema flexibility for both AI consumption and customer-facing queries.
Set standards for data quality, validation, and lineage- every dataset that reaches an AI model or a customer dashboard must have a defined owner, a quality score, and a freshness SLA.
Integration Engineering and Data Acquisition:
Work with the Integrations team to ensure the Apache Camel-based integration framework is scalable, maintainable, and well-documented- covering ERP systems (SAP, Oracle), distributor APIs, CRM platforms, and market data feeds.
Set architecture standards for BOT development and web scraping- the integrations team builds automated data acquisition pipelines that extract component availability, pricing, and supplier data from distributor portals and public sources at scale; these must be reliable, stealthy, and rate-limit-aware.
Define API integration patterns: authentication strategies (OAuth2 API keys, mTLS), rate-limit handling, retry logic, idempotency, and structured error contracts- enforced consistently across all third-party integrations.
Ensure all integration and scraping pipelines are observable end-to-end: every message flow has a trace, every failure has an alert, and every data acquisition job has a freshness SLA that is monitored.
Team Leadership and People Development:
Coach and manage a team of 10+ engineers across four streams- assign work, review code, set goals, run 1:1s, deliver feedback, hire, and onboard.
Partner with stream leads (Backend, Frontend, Data Engineering, Integrations) to set quarterly goals, run effective sprint ceremonies, and maintain delivery predictability without bureaucratic overhead.
Create an engineering culture defined by craft, ownership, and psychological safety- where engineers raise issues early and fix problems permanently.
Run a lean Agile process that actually helps: shape work with Product, estimate with honesty, run effective standups and retros, and keep delivery cadence visible to the CPTO.
Own hiring for the engineering function- define role profiles, run structured interviews, and build a team that raises the average with every hire.
Level up engineers through deliberate pairing, code review, and architecture exposure- not just annual performance cycles.
Requirements:
You are a hands-on engineer first.
You write production code every week.
You review PRs not just for logic but for performance, security, and maintainability.
You would not ask your team to do something you could not do yourself.
11+ years in software engineering with a strong track record of shipping production systems- not prototypes, not internal tools, but systems that served real customers under real load.
At least 4-5 years managing and mentoring engineers while remaining deeply technical.
You have grown engineers, not just assigned them tickets.
Expert-level system design capability.
You have designed systems under real constraints- high throughput, low latency, partial failure, and evolving requirements.
You produce clear, reasoned architecture documentation and can defend every decision under scrutiny.
Deep microservices architecture experience.
You have designed, decomposed, and operated microservices at production scale- you understand service boundaries, distributed consistency, API contracts, versioning, and the real cost of over-decomposition.
Event-driven architecture in production.
You have owned Kafka at scale- topic design, partitioning, consumer group management, Schema Registry, Debezium CDC, and debugging consumer lag and ordering issues in live systems.
Event-driven is not a pattern you have read about; it is how you build.
Deep Node.js expertise- event loop, async patterns, performance profiling, and production debugging.
You have shipped high-throughput Node.js services and know where the runtime bites you.
React / Next.js proficiency- you can meaningfully review frontend architecture, SSR vs CSR trade-offs, state management, and component design.
You do not need to be a CSS expert, but you cannot be blind to the frontend.
MongoDB in production- schema design, indexing strategy, aggregation pipelines, sharding trade-offs, and performance tuning under load.
AWS architecture experience- ECS/EKS, Lambda, RDS, S3 networking (VPC, security groups), and cost optimisation.
You can design a new service's infrastructure from scratch without a dedicated DevOps team.
DevOps and SRE practices- CI/CD, IaC (Terraform), Kubernetes, observability (tracing, metrics, logging), on-call rotation, and incident management.
You have owned production reliability personally.
Comfortable being a polyglot. Rust is on the stack.
You do not need to be an expert on day one, but you learn new languages fast and are curious, not resistant.
Apache Camel working knowledge- mediation, routing, and transformation patterns well enough to review integration designs, evaluate BOT scraping architectures, and debug pipeline failures.
Clear, direct commun
Experience
11-15 yrs
More openings worth a look
Recently tracked roles with full details and direct application links.