A dependable path to publication.
Source-specific adapters feed a shared processing flow. Listings are validated, normalized, and checked for duplicates before publication, while raw records remain available for inspection and reprocessing.
From fragmented sources to searchable opportunities, with control over every step in between.

For someone looking for work, an incomplete or outdated listing is a lost opportunity. Sidehustles brings flexible jobs into one place. We engineered the cloud and data systems behind that experience, connecting collection, processing, publication, and operations into a pipeline the team can understand and manage.
A multi-source ingestion system with dedicated workers, validation and duplicate prevention, a PostgreSQL data layer, and search integration. Around it, we built containerized cloud deployments, release-time database checks, queue health monitoring, and an operations dashboard for tracing a listing back to its source.
Source-specific adapters feed a shared processing flow. Listings are validated, normalized, and checked for duplicates before publication, while raw records remain available for inspection and reprocessing.
Background workers separate collection from processing. Queue health checks expose backlogs and stalled progress, giving operators a practical way to diagnose ingestion without treating every failure as an application outage.
Database migrations run before a new web release is promoted. If that step fails, the previous web process keeps serving. The operations dashboard brings run status, validation errors, and source data into one investigation workflow.
The hard part was keeping the whole journey dependable: different source formats, changing pages, growing processing workloads, and a public experience that relies on consistent data.
Collection and processing are decoupled. Data quality is checked before publication. Cloud releases and operational visibility support the entire lifecycle.
Employer pages, ATS platforms & feeds
Python actors and browser collection for dynamic pages
Ingest through Go APIA database-backed queue retains pending work.
Process pending workValidation, normalization and duplicate checks. Invalid records stay identifiable.
Publish valid recordsPrimary writes and read-replica query roles.
Query the catalogCatalog filters, hybrid retrieval and results for the public experience.
Workers and the API turn listing and query text into vectors.
Workers index listings; the API retrieves by meaning and keywords.
Next.js compares raw and normalized records and investigates Apify runs.
Dokku runs migrations before web promotion. Docker keeps services reproducible.
Queue status, invalid records and worker progress. Separate environments for validating changes.
The ingestion API preserves raw records in PostgreSQL. Workers process pending work, validate and normalize listings, then publish them to the catalog. OpenAI and Qdrant extend that journey with embeddings and indexing for hybrid discovery.
The Next.js dashboard inspects data and runs without changing listings. Queue health checks expose processing progress. Before a web release is promoted, Dokku runs database migrations; if that step fails, the previous web process keeps serving.
The platforms and engineering tools connecting ingestion, discovery, and cloud operations.
Cloud hosting and managed PostgreSQL infrastructure.
Service deployment with release-time migration checks.
Containerized services and repeatable deployment artifacts.
Raw and processed listing storage, with primary and read-replica roles.
Ingestion APIs, background processing, and queue health checks.
Source adapters, extraction, validation, and normalization.
Actor execution and visibility into scraping runs.
Browser-based collection for dynamic source pages.
Vector retrieval for semantic and related-job discovery.
Text embeddings used by the semantic search pipeline.
An internal dashboard for comparing raw and normalized data and investigating runs.
Follow each listing from its source to the public experience. Define where collection ends, processing begins, and publication becomes safe, including how rejected records remain inspectable.
Keep source adapters, processing workers, and user-facing services distinct. Preserve raw input, normalize a consistent record, and make queue progress observable.
Treat migrations, environment separation, release verification, and diagnostics as part of the platform. Give operators a way to investigate failures without changing production listing data.
Built integrations for ATS sources including Workday and iCIMS, alongside external feeds. Rate limiting, structured extraction, and source identifiers make collection easier to extend and diagnose.
Implemented duplicate checks and validation, retained raw payloads, and normalized fields such as location and salary. Invalid records remain distinct from published listings.
Aligned hub counts with search filtering and connected keyword and semantic retrieval. Discovery depends on both relevant results and a consistent picture of what is available.
Added queue health checks and a read-only inspection workflow for source runs, normalized records, and validation errors. Release checks protect the web process when a migration fails.
Start with a technical conversation about the infrastructure issues creating risk, instability, or friction for your team.
Talk to an expert