Live opening · Posted 4 days ago

Director - Infra Engineering - Platform Security and Lifecycle Management

American Express · Phoenix, AZ, United States | New York, NY, United States | Palo Alto, CA, United States
Oracle Hybrid
You are 4 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 4 days ago
CompanyAmerican Express
LocationPhoenix, AZ, United States | New York, NY, United States | Palo Alto, CA, United States
Work modeHybrid
SourceOracle
Listed4 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
0 min from Oracle publishing this role to us finding it
25 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
21,317 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

American Express Enterprise Cloud team is looking for innovators to help us build world-class applications, Cloud platforms and infrastructure supported by integrated CICD, Observability and security capabilities.
The Director Infrastructure Engineering – Platform Security and Lifecycle Management is responsible for leading the strategy, governance, and execution of Platform as a Service and Middleware services lifecycle management across Private Cloud environments. This role leverages Gen AI/Agentic AI to drive fully automated platform upgrades, security posture management, capacity management, and operational resilience to ensure a secure, scalable, highly available, and compliant Private Cloud platform supporting mission-critical business applications.
The Director partners closely with Platform Engineering, Information Security, Architecture, Infrastructure, SRE, DevOps, and Application Development teams to drive platform modernization, automate operations, reduce technology risk, and enable cloud-native adoption at scale.
Bachelor’s degree in Computer Science, Engineering, or related field (Master’s preferred).
8+ years of experience in Platform Engineering & Operations, Site Reliability Engineering (SRE), Platform lifecycle management with a proven track record of leading teams in managing large-scale cloud infrastructure with a focus on automation, reliability and resilience.
Deep hands-on experience with any Kubernetes platform(multi-cloud preferred).
Experience building end to end platform upgrade and fleet management automation leveraging Gen AI / Agentic AI
Strong experience with:
Infrastructure as Code (Terraform, CloudFormation, ARM)
Infrastructure automation tools like Ansible
Container platforms (OpenShift/Kubernetes)
Monitoring tools (Prometheus, OTEL, LOKI)
CI/CD pipelines (Jenkins, GitHub Actions)
Open source based messaging, caching, and database technologies like Kafka, Redis, Elastic
Strong understanding of cloud networking, security, and architecture.
Experience managing large-scale, mission-critical production environments.
Relevant certifications preferred
Experience with DevOps practices and methodologies, including CI/CD pipelines, configuration management, and infrastructure as code.
Experience with observability tools such as Prometheus, Splunk, ELK, Dynatrace.
Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
Excellent communication and leadership skills, with the ability to effectively collaborate with cross-functional teams and influence decision-making at all levels of the organization.
Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.
Private Cloud Platform Upgrade Strategy & Modernization
Define and execute the enterprise strategy and roadmap for Enterprise PaaS platform using Redhat Openshift and Data Middleware platform upgrades and lifecycle management.
Lead major version upgrades, cluster modernization, and infrastructure refresh initiatives across development, test, and production environments.
Establish standards, reference architectures, and best practices for platform lifecycle and security management.
Leverage Gen AI/Agentic AI to drive a fully automated pipeline for platform provisioning, upgrades, patching, and configuration management
Upgrades & Release Management
Own end-to-end upgrade planning, governance, risk assessment, and execution.
Establish upgrade readiness processes, validation frameworks, rollback strategies, and post-upgrade monitoring.
Coordinate with application teams to ensure platform compatibility and minimize business disruption during upgrades.
Manage lifecycle risks associated with OpenShift, Kubernetes, operating systems, middleware, and supporting infrastructure.
Track and report platform currency and technology lifecycle compliance metrics.
Security & Compliance
Establish and maintain a strong security posture across clusters and supporting infrastructure.
Lead vulnerability management, container image security, platform hardening, patch management, and remediation efforts.
Ensure compliance with enterprise security policies, regulatory requirements, and industry standards.
Partner with Information Security and Risk teams to manage security assessments, audits, and remediation activities.
Manage security exceptions and risk acceptance processes while driving long-term remediation strategies.
Capacity & Performance Management
Own capacity planning and forecasting for Platform, including compute, memory, storage, and network resources.
Develop predictive capacity models to support business growth and application onboarding.
Establish monitoring and reporting processes for cluster utilization, performance, and scalability.
Optimize infrastructure consumption and platform efficiency while maintaining service-level objectives.
Lead resource optimization initiatives to improve workload density and reduce infrastructure costs.
Leadership & Team Management
Build and lead high-performing cloud native DevOps engineers, Kubernetes administrators, security specialists, and capacity planners.
Foster a culture of operational excellence, continuous learning, innovation, and accountability.
Mentor leaders and technical experts within the organization.
Drive workforce planning, succession planning, and talent development initiatives.
Stakeholder & Vendor Management
Serve as the senior technology leader for Private Cloud platform services.
Partner with application development, architecture, security, infrastructure, and business leaders to align platform capabilities with organizational priorities.
Manage relationships with key vendors, managed service providers, and strategic partners.
Present platform health, upgrade status, security posture, risks, and capacity forecasts to executive leadership.

Work arrangement
Hybrid

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App