Live opening · Posted 6 days ago
At a glance
The key details from the original listing.
Your early-applicant advantage
Live timing from JobBeeper.
About the role
Description supplied by the original job listing.
Company
Canaria is a technology startup providing high-quality, large-scale labor market data to the B2B market. Our mission is to power the next generation of solutions in recruitment, HR analytics, workforce planning, lead generation, healthcare, academic research, and finance by delivering the most comprehensive and accurate job market data available. Our advanced data mining, computing optimizations, and state-of-the-art NLP techniques let us process job market data at scale: over 1 billion unique job postings drawn from 200,000+ data sources, with 100+ fields per record. Our clients use this data to understand market trends, identify skill gaps, and connect people with the right opportunities.
What You Will Do
Write production Python daily. This is a Python-first role.
Own the design and implementation of large-scale Python data pipelines (ingestion, parsing, cleaning, semantic deduplication, enrichment) over billions of job postings.
Extend our in-house distributed crawling system: our own scheduling, queueing, proxy and session management, rendering, and anti-blocking layers, built and tuned by us rather than assembled from off-the-shelf frameworks. You will be adding new source integrations and making the core faster and harder to block.
Design and operate the databases behind that pipeline across PostgreSQL, MongoDB, Redis and Aerospike, at a scale where query plans and storage layout actually matter.
Deploy your Python services with Docker on AWS/GCP and keep them running at scale.
Ensure system quality through stress, integration, and end-to-end testing.
Write clean, maintainable, well-documented code.
Who You Are (Required)
3+ years in a paid role where Python was your primary language.
3+ of those years on production backend systems that were data-intensive (high-volume ETL, scraping, streaming, or large-scale batch processing).
Completed Bachelor's or Master's degree in Computer Science, Computer Engineering, Software Engineering, or a closely related computing field. Non-computing engineering degrees (mechanical, civil, mechatronics, EEE), statistics/mathematics degrees, associate degrees, and in-progress degrees do not meet this requirement.
Comfortable in a Unix environment for both scripting and application development.
Production experience with at least two named database systems, one relational (PostgreSQL, MySQL) and one NoSQL (MongoDB, Redis, Aerospike, ClickHouse).
Hands-on experience with at least one major cloud platform (AWS, Google Cloud, or Azure).
Comfortable containerizing and deploying applications with Docker.
Nice to Have
Kubernetes in production.
High-throughput key-value stores beyond the basics (Aerospike tuning, Redis clustering).
RESTful API design and optimization.
CI/CD pipelines and automated testing frameworks.
Monitoring and logging tools (e.g. Prometheus, Grafana).
Big data warehouses (BigQuery, Redshift).
Benefits
High-end laptop.
Annual performance bonus.
Work with global teams (US, EU).
Founders and colleagues with FAANG and top-institution backgrounds.
Access to powerful on-prem and cloud compute for research.
Professional development and growth opportunities.
Hiring Process
First call about the role & expectations (15 mins)
Python coding interview (1 hour)
Technical interview (1 hour)
Work arrangement
Yes
More openings worth a look
Recently tracked roles with full details and direct application links.