Our work

Reliable job data. Infrastructure built to keep it moving.

From fragmented sources to searchable opportunities, with control over every step in between.

The project

For someone looking for work, an incomplete or outdated listing is a lost opportunity. Sidehustles brings flexible jobs into one place. We engineered the cloud and data systems behind that experience, connecting collection, processing, publication, and operations into a pipeline the team can understand and manage.

What we delivered

A multi-source ingestion system with dedicated workers, validation and duplicate prevention, a PostgreSQL data layer, and search integration. Around it, we built containerized cloud deployments, release-time database checks, queue health monitoring, and an operations dashboard for tracing a listing back to its source.

Services

  • Cloud infrastructure & operations
  • Data ingestion & scraping pipelines
  • Data quality & normalization
  • Search infrastructure & integration
  • Deployment & database engineering
  • Operational visibility & diagnostics

Outcomes & value created

The platform, at scale.

Job listings on the platform
800k+
Companies represented
15k+
Catalog figures reported by Sidehustles

A dependable path to publication.

Source-specific adapters feed a shared processing flow. Listings are validated, normalized, and checked for duplicates before publication, while raw records remain available for inspection and reprocessing.

More control as the workload grows.

Background workers separate collection from processing. Queue health checks expose backlogs and stalled progress, giving operators a practical way to diagnose ingestion without treating every failure as an application outage.

Changes with a safer release path.

Database migrations run before a new web release is promoted. If that step fails, the previous web process keeps serving. The operations dashboard brings run status, validation errors, and source data into one investigation workflow.

Objectives & challenges

The hard part was keeping the whole journey dependable: different source formats, changing pages, growing processing workloads, and a public experience that relies on consistent data.

  • Ingest employer pages, ATS platforms, and external feeds with different structures.
  • Preserve source data while producing consistent, searchable listings.
  • Identify duplicates, invalid records, and processing bottlenecks.
  • Evolve services and schemas while keeping the public platform available.

From source data to dependable discovery.

Collection and processing are decoupled. Data quality is checked before publication. Cloud releases and operational visibility support the entire lifecycle.

Cloud & data architectureComponents and responsibilities

External sources

Employer pages, ATS platforms & feeds

Collection

Python actors and browser collection for dynamic pages

Ingest through Go API

DigitalOcean platform

Containerized services · Managed PostgreSQL
  1. Preserve the source

    Raw records + processing state

    A database-backed queue retains pending work.

    Process pending work
  2. Process asynchronously

    Dedicated worker pool

    Validation, normalization and duplicate checks. Invalid records stay identifiable.

    Publish valid records
  3. Publish the catalog

    Normalized job listings

    Primary writes and read-replica query roles.

    Query the catalog
  4. Serve discovery

    Search & category API

    Catalog filters, hybrid retrieval and results for the public experience.

DokkuDocker

External search services

Workers & API ↔ embeddings
OpenAIText embeddings

Workers and the API turn listing and query text into vectors.

Workers → index · API ↔ search
QdrantHybrid index & retrieval

Workers index listings; the API retrieves by meaning and keywords.

Operations & delivery

Read-only diagnostics

Next.js compares raw and normalized records and investigates Apify runs.

Controlled releases

Dokku runs migrations before web promotion. Docker keeps services reproducible.

Processing health

Queue status, invalid records and worker progress. Separate environments for validating changes.

Arrows follow the listing lifecycle. Labeled connections describe calls to external services. Boundaries group responsibilities without revealing private configuration.

The ingestion API preserves raw records in PostgreSQL. Workers process pending work, validate and normalize listings, then publish them to the catalog. OpenAI and Qdrant extend that journey with embeddings and indexing for hybrid discovery.

The Next.js dashboard inspects data and runs without changing listings. Queue health checks expose processing progress. Before a web release is promoted, Dokku runs database migrations; if that step fails, the previous web process keeps serving.

Cloud, data & search stack

The platforms and engineering tools connecting ingestion, discovery, and cloud operations.

Cloud & delivery

  • DigitalOcean

    Cloud hosting and managed PostgreSQL infrastructure.

  • Dokku

    Service deployment with release-time migration checks.

  • Docker

    Containerized services and repeatable deployment artifacts.

  • PostgreSQL

    Raw and processed listing storage, with primary and read-replica roles.

Ingestion & processing

  • Go

    Ingestion APIs, background processing, and queue health checks.

  • Python

    Source adapters, extraction, validation, and normalization.

  • Apify

    Actor execution and visibility into scraping runs.

  • Chrome / Browserless

    Browser-based collection for dynamic source pages.

Discovery & operations

  • Qdrant

    Vector retrieval for semantic and related-job discovery.

  • OpenAI

    Text embeddings used by the semantic search pipeline.

  • Next.js

    An internal dashboard for comparing raw and normalized data and investigating runs.

Our approach

  1. Map the data lifecycle.

    Follow each listing from its source to the public experience. Define where collection ends, processing begins, and publication becomes safe, including how rejected records remain inspectable.

  2. Separate work that scales differently.

    Keep source adapters, processing workers, and user-facing services distinct. Preserve raw input, normalize a consistent record, and make queue progress observable.

  3. Engineer the operating model.

    Treat migrations, environment separation, release verification, and diagnostics as part of the platform. Give operators a way to investigate failures without changing production listing data.

Key activities & takeaways

Adapters that respect the source.

Built integrations for ATS sources including Workday and iCIMS, alongside external feeds. Rate limiting, structured extraction, and source identifiers make collection easier to extend and diagnose.

Data quality before discovery.

Implemented duplicate checks and validation, retained raw payloads, and normalized fields such as location and salary. Invalid records remain distinct from published listings.

Search that matches the catalog.

Aligned hub counts with search filtering and connected keyword and semantic retrieval. Discovery depends on both relevant results and a consistent picture of what is available.

Visibility beyond a green deployment.

Added queue health checks and a read-only inspection workflow for source runs, normalized records, and validation errors. Release checks protect the web process when a migration fails.

Ready to stabilize and scale your cloud?

Start with a technical conversation about the infrastructure issues creating risk, instability, or friction for your team.

Talk to an expert