Live opening · Posted 7 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
We're looking for a Lead AI Engineer skilled in Spark and AWS Services to become part of the RBQM Production Pod within the program. In this position, you'll construct and sustain data pipelines that drive AI/GenAI applications supporting Risk-Based Quality Management for clinical trials. The primary focus of this role includes RAG document ingestion, vector indexing, and developing data APIs for AI applications.
Responsibilities
Architect and construct RAG document ingestion pipelines (chunking, embedding, vector indexing) to support clinical trial quality data
Establish and oversee vector databases (AWS OpenSearch) to support RAG-driven AI workflows
Create batch and streaming ETL/ELT pipelines from the ground up for unstructured clinical data (PDF, DOCX, clinical reports)
Construct and expose data APIs that AI applications can consume
Enhance chunking strategies, embedding generation, and retrieval performance within RAG architectures
Oversee data quality, lineage, and governance across AI/ML data pipelines
Set up and sustain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)
Partner with Data Scientists and Backend Developers as part of a unified pod team
Requirements
Minimum 7 years of practical, large-scale data engineering experience
Strong background in RAG document ingestion pipelines (chunking, embedding, vector indexing)
Skilled in using AWS OpenSearch as a vector database for RAG workflows
High-level command of Python, along with SQL and Spark SQL
Experience transforming unstructured data (PDF, DOCX) for use in RAG/LLM applications
Working knowledge of AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB
Understanding of Docker-based containerization
Ability to develop custom pipelines from the ground up, going beyond simple configuration of pre-built services
Proficiency in English at a B2+ level
Nice to have
Experience within the pharmaceutical or life sciences sector
Exposure to Snowflake and Pinecone (as an alternative vector database)
Understanding of SageMaker processing jobs
Proficiency with CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure-as-code tools (CDK or Terraform)
Familiarity with clinical data standards (CDISC, ADaM, SDTM)
We offer
CONTINUOUS UPSKILLING, LEARNING & DEVELOPMENT
Diversity of tasks and projects
Assessment center for objective review of competency level
Personal development plan
Mentoring programs and leadership development
Certification and professional development support
Access to learning platforms including more than 2,500 internal courses
English courses taught by certified teachers
CORPORATE BENEFITS
Extra leave days
Referral bonuses
COMPENSATION PACKAGE
Competitive compensation paid in USD
Regular salary and performance reviews
MEDICAL & HEALTHCARE
Private health insurance
Well-being events
WORKING ENVIRONMENT
Recreation areas and kitchens
Tea, coffee and snacks
Sports equipment and game consoles
IT Equipment
Microsoft’s Software Assurance Home Use Program (HUP)
Please note that our Talent Attraction Team reviews applications and CVs submitted in English.
EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.