← Back to work
Drone-Trajectory Data Pipeline
AirflowdbtPostgreSQLDocker

Drone-Trajectory Data Pipeline

Training project

A fully dockerized data-warehouse stack for traffic-trajectory data from swarm drones (pNEUMA): Airflow ingests large CSVs (~87MB each), and dbt builds tested, documented staging/production models for spatial-temporal analysis.

The problem

Traffic-trajectory data captured by swarm drones (the pNEUMA dataset) arrives as large CSV files that are awkward to store, transform, and query.

How it works

  1. 01, Input
    pNEUMA CSVs

    Swarm-drone trajectory files of around 87MB each.

  2. 02, Tool or code
    Airflow ingestion

    Loads the large CSVs into the warehouse.

  3. 03, Tool or code
    PostgreSQL warehouse

    Stores the raw trajectory data.

  4. 04, Tool or code
    dbt models

    Tested, documented staging and production schemas.

  5. 05, Output
    Analysis-ready data

    Queryable data for spatial-temporal analysis of vehicle paths.

InputTool or codeOutputThe whole stack runs in Docker.

What I did

  • Built a fully dockerized data-warehouse stack.
  • Used Airflow to ingest the large trajectory CSVs, around 87MB each.
  • Modeled the data with dbt, including tests, documentation, and staging and production schemas.
  • Made the workflow reproducible for spatial-temporal analysis of vehicle paths.

Results

  • Turned bulky raw CSVs into a queryable, tested warehouse.
  • Documented a reproducible workflow for spatial-temporal analysis.