Live opening · Posted 9 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Data Engineer | ETL, OCR & AI Document Processing
Location: Pakistan
Work Mode: Fully Remote
Employment: Full-Time
Experience: Fresh graduate or up to 2 years
Working Hours: 7:00 PM to 3:00 AM PKT, Monday to Friday
Start Date: Must be available to start next week
Compensation: Based on experience and technical capability
About the Role
Clearprise is hiring a Junior Data Engineer to work full-time with an international AI and software company.
You will help build data pipelines and intelligent document-processing systems that convert scanned PDFs, unstructured documents, APIs, and multi-source operational data into accurate, structured information for AI models, predictive systems, and SaaS applications.
This is a hands-on role for a talented fresh graduate or junior engineer with strong Python, SQL, ETL, API, and data-processing fundamentals.
You do not need extensive corporate experience, but you must be able to demonstrate relevant projects, explain your technical decisions clearly, and work independently in a fast-moving remote environment.
Key ResponsibilitiesETL and Data Pipelines
Build and maintain ETL pipelines
Extract data from APIs, databases, CSV files, PDFs, and other sources
Clean, normalize, map, and transform raw data
Load structured data into relational databases and downstream systems
Develop repeatable and reliable data-processing workflows
Handle missing, duplicate, invalid, and inconsistent data
Monitor pipeline performance and resolve failures
Intelligent Document Processing
Process scanned PDFs, images, invoices, and other unstructured documents
Integrate OCR and document-extraction tools
Convert extracted document content into structured JSON or database records
Validate extracted fields using schemas and business rules
Handle unclear documents, inaccurate extraction, and low-confidence results
Support human-review workflows for uncertain data
Preserve original documents and maintain processing records
Backend and API Integration
Develop data-processing services using Python
Build or integrate REST APIs
Connect external systems, databases, and AI services
Handle authentication, pagination, rate limits, and API failures
Create structured data feeds for AI and predictive models
Support backend functionality for AI-governance and document-processing SaaS products
Database and Data Quality
Work with PostgreSQL, Supabase, or comparable databases
Design tables, schemas, relationships, and validation rules
Write and optimize SQL queries
Maintain data quality, consistency, and traceability
Implement logging, retry handling, and error reporting
Support audit trails and data-governance requirements
Collaboration and Delivery
Work directly with an international founder and technical team
Provide clear progress updates and communicate blockers early
Document pipelines, APIs, transformations, and technical decisions
Use Git and follow version-control workflows
Take ownership of assigned tasks with minimal supervision
Learn new tools and technologies quickly
Mandatory Requirements
Fresh graduate or up to two years of professional experience
Currently based in Pakistan
Strong Python programming fundamentals
Working knowledge of SQL and relational databases
Understanding of ETL concepts and data-pipeline architecture
Hands-on project or professional experience collecting, cleaning, transforming, and storing data
Experience working with REST APIs and JSON
Exposure to OCR, document processing, AI integrations, or unstructured-data extraction
Understanding of data validation, error handling, and logging
Experience using Git and version control
Strong debugging and problem-solving ability
Professional English communication
Previous experience working directly with an international client, employer, or technical team
Ability to work from 7:00 PM to 3:00 AM PKT, Monday to Friday
Ability to start next week
Ability to work independently in a remote environment
Strongly Preferred
Recent degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or a related field
Experience with Pandas, NumPy, or similar Python libraries
Experience with FastAPI or another Python backend framework
Experience with PostgreSQL or Supabase
Experience extracting structured data from PDFs and scanned documents
Familiarity with OCR services or libraries
Experience integrating LLMs or AI APIs into data workflows
Familiarity with data schemas, validation frameworks, and confidence scoring
Experience with Docker
Exposure to AWS, Azure, GCP, or another cloud platform
Experience deploying a data pipeline or backend service
Relevant GitHub projects or technical portfolio
Ideal Candidate
You are an early-career engineer with strong technical fundamentals and evidence that you can build practical data solutions.
You can explain where data came from, how you cleaned and transformed it, where it was stored, how errors were handled, and how the final output was used.
You understand that AI and OCR results must be validated before entering a production database. You are curious, dependable, comfortable asking questions, and capable of taking ownership of clearly assigned work.
Employment Details
Full-time and fully remote
Eight hours per day, Monday to Friday
Working hours from 7:00 PM to 3:00 AM PKT
Direct collaboration with an international client
Long-term opportunity
Required start date: next week
Work arrangement
No
More openings worth a look
Recently tracked roles with full details and direct application links.