Open source AV data infrastructure

Data infrastructure
for India’s roads.

An open annotated dataset of Indian roads: 645,714 frames and 3.67M tracked detections from Delhi NCR, free on Hugging Face.

Annotated frames
600K+
HuggingFace downloads
~550/mo
Detection classes
12
Commercial competitors
0

Platform coverage

Most AV datasets come from Western roads. We built one from Delhi’s.

Platform

Four layers from sensor to shipping model.

Together these layers take an AV company from “we want to enter India” to “we are deploying on Indian roads”, without building any of the data infrastructure themselves.

01Live

Road dataset

645,714 annotated frames from Indian roads.

Dashcam capture from the Delhi NCR region, paired with per-clip GPS. 3.67M 2D bounding boxes with track IDs and 1.29M semantic segmentation masks, across a 12-class detection taxonomy built for Indian roads.

  • 3.67M tracked detections
  • 12 detection classes
  • 1.29M segmentation masks
  • Open source on HuggingFace
02

RL environments

Coming soon

Simulation calibrated against real Indian traffic.

Gym-compatible environments built from observed distributions. Train agents on unstructured roads, not synthetic Western intersections.

  • OpenAI Gym compatible
  • Multi-agent
  • Real-world calibrated
03

Benchmarks

Coming soon

Evaluation built for mixed traffic.

Purpose-built suites for unstructured roads. Leaderboard and continuous evaluation against Indian driving conditions.

  • Mixed traffic scenarios
  • Long-tail coverage
  • Public leaderboard
04

Fine-tuning and compute

Planned

Managed GPU pipelines for model training.

Architecture-agnostic training on our data and environments. Production export included.

  • Managed GPU
  • Experiment tracking
  • Architecture agnostic

Open source core. Enterprise tiers available.

Talk to us

RL environmentsComing soon

Where perception meets decision-making.

Gym-compatible environments calibrated against real Indian traffic distributions. Every scenario built from observed data.

Vehicle agents
Independently controlled vehicles following their own routes
Pedestrians
Crossing on and off the marked path
Auto-rickshaw
Three-wheeler traffic, a class absent from Western datasets
Detection range
Concentric sensor range rings around the ego vehicle
train.py
import thirdeyelabs
 
env = thirdeyelabs.make("UrbanIntersection-v2")
obs, info = env.reset()
 
for step in range(10_000):
action = agent.predict(obs)
obs, reward, done, _, info = env.step(action)
if done:
obs, info = env.reset()
 
results = thirdeyelabs.evaluate(
agent, suite="unstructured-v1", episodes=500
)
print(f"Success: {results.success_rate:.1%}")

pip install thirdeyelabs · Apache 2.0

Environment registry

Planned reinforcement learning environments, with concurrent agent count and difficulty
ScenarioAgentsDifficulty
Urban intersection12-40Hard
Highway merge8-25Medium
Market road20-60Extreme
Night monsoon10-30Extreme
Rural highway4-15Medium
Roundabout15-35Hard

Pipeline

From road to production.

A vertically integrated pipeline. Every step informed by the one before it.

  1. 01

    Capture

    FleetRaw data

    Dashcam-equipped cabs across the Delhi NCR region. Each clip is captured at 1080p30 with a synchronised GPS track, then sampled to keyframes.

  2. 02

    Annotate

    Raw dataLabels

    Machine labelling with models fine-tuned on Indian road data: detection, tracking, semantic segmentation and scene classification. 3.67M detections across 12 classes, exported in BDD100K format.

  3. 03

    Simulate

    LabelsEnvironments

    Calibrated scenarios replayed in RL environments. Your agent trains on real Indian traffic distributions, not synthetic Western approximations.

  4. 04

    Ship

    TrainingProduction

    Evaluate on our benchmarks, fine-tune on our data, deploy with confidence. Install our package, integrate, and begin training in days.

Contact

Let’s talk about your use case.

Whether you need the open dataset, custom annotations, or a fully managed pipeline, we will have a proposal within one business day.

  • Open source dataset, free to start
  • Enterprise SLAs available
  • Open dataset under CC BY 4.0

Required