From raw multimodal data to training-ready datasets. In one pipeline, at any scale.

Multimodal data engine built for AI.

pip install daft
Amazon logo
Atoms logo
Essential logo
TogetherAI logo
Bytedance logo

Multimodal-native

Process video, images, audio, and sensor data alongside structured metadata in a single dataframe.

CPU and GPU in one pipeline

Run GPU inference/embeddings alongside CPU decode and filter in one pipeline. Daft handles the scheduling and batching, no glue code required.

Python dataframe API

Same operations you use in Pandas or Spark: filter, transform, aggregate, write. No new framework to learn.

Why Daft

  • Open-source data engine (Apache 2.0) with more than 5k GitHub stars, in production at companies like Amazon and Essential AI.
  • Bring your own models, storage, and infrastructure. Run pipelines exactly where and how you want.
  • If you know Pandas or Spark, you know Daft. Load a dataset, filter, transform, and export to training format in a few lines.
View Documentation

Use Cases

AI Search

import daft
from daft.functions import file, prompt, embed_text
from pydantic import BaseModel

class Classifer(BaseModel):
    title: str
    summary: str
    keywords: list[str]

df = (
    daft.from_glob_path("s3://org-data/docs/*.pdf")
    .with_column("metadata", 
        prompt(
            messages=file(daft.col("path")),
            return_format=Classifer, 
            model="gpt-5-mini", 
            provider="openai",
        ),
    )
    .with_column("abstract_embedding", 
        embed_text(
            daft.col("metadata")["summary"], 
            model="text-embedding-3-large"
            provider="openai",
        )
    )
)

# Write to Turbopuffer
df.write_turbopuffer(
    namespace="org-data",
    distance_metric="cosine_distance",
    region="us-west-2",
)
  • This example shows how using LLMs and embedding models, Daft extracts metadata, generates vectors, and writes them to a vector database.
[1]

Native model operators

Embeddings, LLM extraction, and structured outputs as first-class operations. Plug in models from OpenAI, Hugging Face, or your own.

[2]

Multimodal column types

Images, video, audio, text, and embeddings as native column types. Decode, transform, and filter them like any other column.

[3]

Local to production consistency

Define pipelines once. Run them on your laptop or scale across a cluster. Same code, no rewrites.

[4]

Managed UDF runtime

Automatic batching, retries, and error handling for model UDFs. Zero-copy execution powered by Apache Arrow.

[5]

Lower memory footprint

Run the same queries with 5x less memory than alternatives. Jobs that would OOM on Spark or Pandas just work.

[6]

Built in Rust for speed

Daft's core is written in Rust. Decode video, run transforms, and join multimodal data at TB scale without paying Python overhead.

trusted by

“Daft was incredible at large volumes of abnormally shaped workloads - I pointed it at 16,000 small Parquet files in a self-hosted S3 service and it just worked! It's the data engine built for the cloud and AI workloads.”

Tony WangData @ Anthropic, PhD @ Stanford
1/5
5789
stars on github
20x
faster start time
10s
PBs processed daily

Get updates, contribute code, or say hi.

Eventual Blog

Subscribe for the latest updates from the Eventual team.

GitHub Discussions

join

Slack Community

join