Live opening · Posted 21 hours ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Education & Assessment Platforms
About Pratham International
Pratham International is a US 501(c)(3) established to support learning innovations for children and youth in contexts across the globe. We enable the adaptation and spread of proven and scalable education solutions pioneered in India by Pratham Education Foundation to contexts across the global south through programming and technical assistance tailored to local realities. Today, Pratham innovations reach millions of children annually through partnerships across ~30 countries.
Through the Pratham Shah PraDigi Innovation Centre (PraDigi), Pratham International also seeks to create technology innovations for low-resource communities, and support models aimed at learning for school, work, and life for children and youth. We work on several cutting-edge projects in the domain of Generative AI, Automatic Speech Recognition, and Natural Language Processing.
The team's innovative technology products include the PadhAI: Reading Assessment App, which uses advanced speech recognition to evaluate children's reading fluency in 8 Indian languages and 6 global languages; Anytime Testing Machine (ATM), an AI-enabled question generation and grading system that allows learners to be assessed on any topic, anytime; and a suite of AI-powered chatbots designed for diverse learning contexts: from early childhood care, to youth employment readiness, and OCR-based analysis tools that digitize and assess student work at scale.
At Pratham International, we build for the long term. Whether it's ensuring products are truly accessible across rural India and the global south through languages, building datasets to ensure context and accuracy for products, or gathering feedback from the field. Backed by partners like Google.org, Anthropic, and the Gates Foundation, we are committed to building digital public goods that ensure every child is in school and learning well.
Role Overview
About the work
We design and deliver AI systems that governments and large education programmes can actually run: curriculum-aligned item generation, paper assembly, handwriting OCR, rubric grading, and teacher-in-the-loop review. The work is already live as practice assessment in India and is being built as a national exam platform in Rwanda, with further country programmes behind it.
We are hiring a Solution Architect who sits between discovery and delivery. You will turn messy workshop notes and ministry requirements into architecture a team can build, secure, cost, and operate — RAG and multi-agent systems, LLM-as-judge pipelines, and the cloud/LLMOps layer underneath.
You will
Own solution design from first workshop to a delivery-ready blueprint (services, data stores, trust boundaries, cost, ops).
Architect LLM pipelines: generation, retrieval over curriculum knowledge graphs, standards/duplication/quality judges, constraint-based paper assembly.
Design human + AI workflows with audit, segregation of duties, cycle-pinned config, and no student PII in model calls.
Choose and govern models, prompts, eval gates, and fallbacks (including OCR: Azure primary, secondary fallback).
Shape the Azure/AWS platform: identity, networking, LLMOps/MLOps, observability, and sovereign-hosting constraints.
Run client workshops, write POVs and architecture packs, and stay accountable until the system is in production.
About the work
Pratham International programmes sit at the intersection of pedagogy, public systems, and production AI. The pattern is consistent across countries; the constraints change.
Anytime Testing Machine (ATM) — India, practice and feedback End-to-end assessment for real classroom conditions: generate curriculum-aligned items, digitise handwritten answers (often photographed on a phone), grade against rubrics, and return feedback a teacher reviews before it reaches the learner. Built for high pupil–teacher ratios, mixed language (e.g. Hindi + English), and uneven connectivity. Pilots have already reached thousands of learners, including Second Chance (women returning for Grade 10). Model quality is measured against expert golden sets, not vendor slides.
National assessment automation A government-grade exam platform. Phase 1 is item authoring and print-ready paper composition before the exam window. Phase 2 is script intake, OCR at national volume, AI-assisted grading with human confirmation on the pilot, results export, and monitoring/evaluation reporting. Subjects span core and elective streams across the curriculum. Design rules that do not exist in a typical SaaS product: cycle-pinned configuration, embargoed papers, hash-chained audit trails, and integration with national student information and assessment management systems.
Curriculum knowledge graphs Structured retrieval corpora (concepts + worked examples) that ground generation and judging. This is the standing public-knowledge exception in an otherwise closed exam system, and the seed of a broader education DPI conversation.
AI in TaRL and teacher support AI tools that help teachers group and teach by learning level, with an RCT-shaped evidence bar. Same product family: human agency stays in the loop; the model is a capacity multiplier.
Country programmes and the three-year direction Adapt the same engine to new ministries, languages, and exam rules (Kenya and other Global South programmes are in scope). The longer bet is a learning and credentialing engine that recognises competence beyond a single syllabus.
This is not a chatbot brief. It is constrained generation, retrieval, judging, document AI, workflow, and public-sector operations.
Why this role exists
Most enterprise AI dies between the workshop and production. We need one person who owns that gap: someone who can sit with NESA, C4IR, Pratham pedagogy, and engineering in the same week, and leave behind a blueprint the delivery team can execute without inventing policy.
The target profile is a practising AI/ML Solution Architect — not a pre-sales slide owner, and not an infra-only platform engineer. The person should have done all three of these:
Designed GenAI / NLP solutions, PoCs, and client workshops, and can talk to research-grade issues (eval, agent safety, prompt-injection, model-specific behaviour).
Owned the AI platform: LLMOps, MLOps, DevSecOps, SRE habits, cloud migration, and FinOps.
Taken messy discovery and turned it into Azure-centred RAG and multi-agent designs that a security team will accept.
Key Responsibilities
Discovery → architecture
Lead solution workshops with programme, ministry, and engineering stakeholders.
Convert half-formed requirements into service cuts, data stores, sequence diagrams, decision gates, and open-question lists with owners.
Write architecture decision records. When the docs disagree (they will), force a decision instead of coding both.
AI system design
Design LLM jobs as contracts: inputs, JSON schema, temperature, persist prompt hash / model / tokens / latency.
Architect generation grounded in structured curriculum blueprints and knowledge-graph retrieval, not "ask the model for a maths paper."
Design LLM-as-judge banks (cognitive-level, relevance, difficulty, language, fairness, standards, duplication, blueprint compliance) with thresholds, one-repair rules, and golden-dataset release gates.
Design hybrid assembly: constraint solver first, model only on ties and layout exceptions.
Design document AI: primary OCR provider, secondary fallback provider, confidence routing, de-identification before any grading call.
Specify eval: inter-rater agreement, rubric-band exact match, first-pass item acceptance, cost per paper cycle.
Platform, trust, cost
Own the reference architecture on the chosen cloud platform(s): identity, RBAC + segregation of duties, private networking, key management, object storage for scans, vector index, data warehouse for monitoring/evaluation reporting.
Draw trust boundaries. Student PII never crosses the model gateway. Prompts are curriculum + knowledge-graph + blueprint content only.
Put FinOps on the design: token budgets, OCR cost caps, degradation behaviour when a cap is hit, model routing.
Specify LLMOps/MLOps: prompt/version registry, eval regression as a release gate, observability, rollback.
Design for sovereign hosting and exam-cycle operations (embargo, watermarking, append-only audit, cycle-pinned configuration).
Delivery and thought leadership
Stay with the build through the first production cycle. Architecture that cannot be implemented is unfinished work.
Produce POVs, sequence packs, and workshop artefacts that can be reused in the next country.
Mentor tech leads on where the model should stop and the workflow should start.
Success in the first two quarters
First 30 days — Orient
Complete architecture walkthroughs of every live/in-flight system with the current engineering leads; document current state (services, data stores, trust boundaries) as-is, not as-designed
Meet every active stakeholder group once: engineering, programme/domain leads, and any external partner or government-side counterpart
Inventory all open technical decisions and disagreements in e
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.