Live opening · Posted 5 days ago

GEAN AI Data Engineer

Virtusa · Hyderabad, Telangana, India (On-site)
Linkedin No
You are 5 days behind. JobBeeper subscribers saw this role while it was still new.

At a glance

The key details from the original listing.

Posted 5 days ago
CompanyVirtusa
LocationHyderabad, Telangana, India (On-site)
Work modeNo
SourceLinkedin
Listed5 days ago

Your early-applicant advantage

Live timing from JobBeeper.

Live data
995 min from Linkedin publishing this role to us finding it
195 min median time from a role going live to a subscriber being told
6 hours subscribers had this role before this page existed
16,411 roles found in the last 24 hours — the newest are not on this site yet
Start your free trial →

About the role

Description supplied by the original job listing.

Job Requirements
Role summary
Create high-quality, customer-specific synthetic data and own RAG / knowledge pipelines so each deployment of
CCAI, voice, and chat can be configured, grounded, demonstrated, and validated without using real customer PII.
You design generation and ingestion pipelines and load data into the correct GCP and AWS services.
What success looks like
Each customer engagement has a documented synthetic dataset covering the channels in scope
Each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable,
and evaluated.
Data and retrieval quality are good enough for configuration, evaluation, and stakeholder demos, and
safe enough for isolation and compliance expectations.
Generation and indexing are parameterized and repeatable, not a one-off manual copy-paste per
customer.
Key responsibilities
Analyze each customer’s domain: intents, entities, knowledge topics, document types, languages, tone,
and edge cases.
Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.
Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structured
entity values.
Schedule and document index refresh processes when customer knowledge changes.
Use appropriate techniques while
documenting parameters and limitations.
Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and source
corpora.
Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized per
customer.
Partner with the Conversational Platform Specialist so loaded data and indexes actually drive the
deployed experience.
Partner with DevOps so pipeline jobs, stores, and secrets are automated and isolated per customer.
Required Qualifications
4+ years in data engineering, conversation design operations, applied NLP data work, or knowledge-
pipeline engineering.
Working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-
grounded knowledge data.
Strong judgment on synthetic-data quality, retrieval quality, and privacy safety.
Preferred Qualifications
LLM-assisted synthetic data generation in a production or implementation setting.
Familiarity with BigQuery, S3, and document stores used as knowledge sources.
Multilingual data generation or evaluation experience.
Work Experience
7-10Years

Work arrangement
No

Get JobBeeper Mobile App

Never miss a job opening! Get instant job alerts on your phone.

Subscribers see fresh openings within minutes. Download the JobBeeper App on Google Play to get real-time push notifications and apply before anyone else.

⚡ Instant Push Alerts 🎯 Tailored Filters 🚀 Direct Employer Links
GET IT ON Google Play

More openings worth a look

Recently tracked roles with full details and direct application links.

6 roles
Good roles move before most people even see them. Tell JobBeeper what you want and get fresh matches delivered in minutes.
Start your free trial →
⚡ Get fresh job alerts 📱 Get App