Live opening · Posted 4 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We are looking for a Cloud
Solution Architect whose primary responsibility is to design, build, manage,
and optimize the enterprise cloud infrastructure that runs our AI
ecosystem.
Job Description:
Role: Cloud Solution
Architect – Enterprise AI Infrastructure & FinOps
Experience: 12–18 years total, with at least 6 years in cloud
architecture and at least 3 years architecting production AI/ML or GenAI
platforms
Domain: Enterprise AI Platform (Pharma / Life Sciences, GxP-aware)
Education: B.E./B.Tech in Computer Science, Information Technology,
Information Security, Electronics & Communication, or related Engineering
disciplines.
Location: Ahmedabad/Mumbai
Role Summary
The individual hired as Cloud
Solution Architect will be responsible for cloud foundation end to end: landing
zones, networking, compute, storage, identity, security, Kubernetes platforms,
GPU capacity, infrastructure-as-code, monitoring, resilience, and day-to-day
infrastructure operations.
The role will ensure that AI and
GenAI workloads operate securely, reliably, and at scale while driving cloud
governance, operational excellence, and FinOps practices to maintain
transparency, control, and optimization of cloud and AI spending.
This is a hands-on infrastructure
leadership role. Application and AI engineering teams build agents and
workflows; you make sure the platform they run on is secure, stable, scalable,
compliant, and cost-efficient.
Key Responsibilities
1. Cloud Infrastructure
Architecture
Own the enterprise cloud infrastructure
architecture on Azure, AWS, and/or GCP, including multi-cloud and
hybrid designs.
Design and maintain enterprise landing zones:
management group and account/subscription structure, policies, guardrails,
naming and tagging standards, and environment separation (dev, test,
validation, production).
Architect network infrastructure:
hub-and-spoke or Virtual WAN topologies, VNets/VPCs, subnets, private
endpoints, DNS, load balancers, application gateways, WAF,
ExpressRoute/Direct Connect, site-to-site VPN, and hybrid connectivity to
data centers and manufacturing sites.
Design compute platforms across VMs, VM scale sets,
containers, Kubernetes (AKS, EKS, GKE), serverless (Functions,
Lambda, Cloud Run), and managed PaaS services.
Architect storage and database infrastructure:
object storage, file shares, managed disks, backup vaults, and managed
databases (SQL, PostgreSQL, Cosmos DB, DynamoDB, Redis).
Design for high availability, multi-region
resilience, disaster recovery, and business continuity with defined
RTO and RPO targets.
Maintain architecture documentation, reference
designs, and architecture decision records (ADRs).
2. Cloud Infrastructure
Management and Operations
Own the day-to-day health, stability, and
performance of cloud infrastructure supporting AI and enterprise
workloads.
Lead infrastructure provisioning, configuration,
patching, upgrades, and lifecycle management for VMs, Kubernetes clusters,
networking, and platform services.
Manage Kubernetes platform operations:
cluster upgrades, node pools (CPU and GPU), ingress, service mesh,
autoscaling, namespaces, and multi-tenancy for AI teams.
Define and track infrastructure SLOs, SLAs, and
capacity plans, and lead major incident response, root cause analysis,
and problem management.
Implement backup, restore, and DR testing on
a regular schedule and maintain runbooks.
Establish ITIL-aligned processes for incident,
change, problem, and service request management in collaboration with IT
operations and managed service providers.
Manage cloud vendors and managed service partners,
including SLA and performance governance.
Drive operational automation to reduce manual
effort and human error.
3. Infrastructure-as-Code,
Automation and Platform Engineering
Define and enforce infrastructure-as-code
standards using Terraform, Bicep/ARM, Pulumi, or CloudFormation.
Build reusable IaC modules, templates, and
"golden paths" so teams can self-serve compliant infrastructure
quickly.
Implement CI/CD and GitOps for
infrastructure (GitHub Actions, Azure DevOps, GitLab, Argo CD, Flux).
Apply policy-as-code (Azure Policy, AWS
SCPs/Config, OPA/Gatekeeper) to prevent misconfiguration and enforce
security and cost guardrails.
Automate environment provisioning, scaling,
patching, compliance checks, and cleanup.
Build an internal developer platform experience for
AI, data, and application teams.
4. AI Infrastructure Enablement
Design and operate the infrastructure foundation
for AI workloads, covering agent runtimes, workflow engines, data
pipelines, model serving, and vector and search services.
Provision and manage GPU and accelerator
capacity (NVIDIA A100/H100/L40S, Inferentia/Trainium, TPUs) for
inference, fine-tuning, and batch processing, including quota management
and capacity reservations.
Host and secure LLM access infrastructure:
Azure OpenAI, AWS Bedrock, Vertex AI, and self-hosted open-source models
(vLLM, TGI, Triton, KServe, Ray).
Deploy and operate a centralized AI/LLM gateway
for secure model access, routing, rate limiting, caching, logging, and
cost attribution.
Provide infrastructure for agent frameworks and
orchestration tools (e.g., LangGraph, Semantic Kernel, Azure AI Foundry,
Bedrock Agents, Temporal, Airflow, Logic Apps, Step Functions), including
MCP servers and secure tool connectors.
Support data platforms used by pipelines, such as
Databricks, Snowflake, Microsoft Fabric, Synapse, Kafka/Event Hub, and
vector stores (Azure AI Search, OpenSearch, pgvector, Pinecone, Qdrant).
Ensure network isolation, private connectivity, and
secure integration between AI services and enterprise systems (SAP, LIMS,
MES, QMS, document repositories, and on-premises data).
Partner with AI engineering teams on LLMOps/MLOps
infrastructure: model registries, experiment tracking, CI/CD for AI
assets, and environment promotion.
Architect
the agent runtime platform: hosting, lifecycle management, state
and memory, tool execution, session handling, and scaling.
Define
standards for agent frameworks and patterns (e.g., LangGraph, Semantic
Kernel, AutoGen, CrewAI, OpenAI Agents SDK, AWS Bedrock Agents, Azure AI
Foundry Agent Service).
Design multi-agent orchest
Experience
12 - 18 Years
Employment type
P-P8-Probationer-HO Executive
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.