Multimodal Data Engine
Built For AI.
From raw multimodal data to training-ready datasets. In one pipeline, at any scale.
From raw multimodal data to training-ready datasets. In one pipeline, at any scale.
Multimodal-native
Process video, images, audio, and sensor data alongside structured metadata in a single dataframe.
CPU and GPU in one pipeline
Run GPU inference/embeddings alongside CPU decode and filter in one pipeline. Daft handles the scheduling and batching, no glue code required.
Python dataframe API
Same operations you use in Pandas or Spark: filter, transform, aggregate, write. No new framework to learn.
Multimodal-native
Process video, images, audio, and sensor data alongside structured metadata in a single dataframe.
CPU and GPU in one pipeline
Run GPU inference/embeddings alongside CPU decode and filter in one pipeline. Daft handles the scheduling and batching, no glue code required.
Python dataframe API
Same operations you use in Pandas or Spark: filter, transform, aggregate, write. No new framework to learn.
“Daft was incredible at large volumes of abnormally shaped workloads - I pointed it at 16,000 small Parquet files in a self-hosted S3 service and it just worked! It's the data engine built for the cloud and AI workloads.”
Why Daft
- Open-source data engine (Apache 2.0) with more than 5k GitHub stars, in production at companies like Amazon and Essential AI.
- Bring your own models, storage, and infrastructure. Run pipelines exactly where and how you want.
- If you know Pandas or Spark, you know Daft. Load a dataset, filter, transform, and export to training format in a few lines.
Use Cases
AI Search
import daft
from daft.functions import file, prompt, embed_text
from pydantic import BaseModel
class Classifer(BaseModel):
title: str
summary: str
keywords: list[str]
df = (
daft.from_glob_path("s3://org-data/docs/*.pdf")
.with_column("metadata",
prompt(
messages=file(daft.col("path")),
return_format=Classifer,
model="gpt-5-mini",
provider="openai",
),
)
.with_column("abstract_embedding",
embed_text(
daft.col("metadata")["summary"],
model="text-embedding-3-large"
provider="openai",
)
)
)
# Write to Turbopuffer
df.write_turbopuffer(
namespace="org-data",
distance_metric="cosine_distance",
region="us-west-2",
)- This example shows how using LLMs and embedding models, Daft extracts metadata, generates vectors, and writes them to a vector database.
Use Cases
Data Enrichment
import daft
from daft.functions import prompt
from typing import Literal
from pydantic import BaseModel, Field
from pydantic.extra_types import phone_number, email_address
EMAIL_REGEX=r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b"
PHONE_REGEX=r"\b\d{3}-\d{3}-\d{4}\b"
# Given a sample product listing
data = daft.from_pylist([{
"id": "1234",
"listing":
"""
Vintage wooden chair $50.
Email john@email.com
Phone 123-456-7890
""",
}])
# And Desired Metadata Schema
class ProductMetadata(BaseModel):
listing: str
category: Literal["furniture", "electronics"]
price: float = Field(ge=0)
has_pii: list[Literal["email", "phone", "none"]]
# Run the pipelines
df = (
data
# Extract Metadata
.with_column("product_meta",
prompt(
daft.col("listing"),
return_format=ProductMetadata,
model="gpt-5-mini",
provider="openai",
),
)
# Redact PII
.with_column("listing_redacted",
when(
lit("email").is_in(
col("product_meta")["has_pii"]
)
).then(
col("listing").regexp_replace(
EMAIL_REGEX,
"[REDACTED_EMAIL]"
)
).when(
lit("phone").is_in(
col("product_meta")["has_pii"]
)
).then(
col("listing").regexp_replace(
PHONE_REGEX,
"[REDACTED_PHONE]"
)
).otherwise(col("listing"))
)
)
# Save results to S3
df.write_parquet("s3://org-data/safe-listings")
- Pipelines label items, normalize attributes, redact PII, and extract structured fields using schema validation and guardrails.
- Daft handles text, images, and other modalities in one place, producing consistent outputs for downstream systems like search, analytics, and product catalogs.
Use Cases
Multimodal AI ETL
@daft.func
def extract_content(document) -> str:
...
@daft.func
def openai_o3_mini(text, prompt) -> str:
...
df = daft.from_glob_paths("s3://financials/10-k/*.pdf")
df = df.with_column("document", df["url"].download())
df = df.with_column("extracted", extract_content(df["document"]))
df = df.with_column("revenue", openai_o3_mini(df["extracted"], prompt="What is the quarterly revenue?"))
df.write_iceberg(iceberg_table)- This example showcases Daft's AI-powered data pipeline capabilities, transforming multimodal data into actionable intelligence.
- The workflow leverages Daft's Rust-based file downloader (40x faster than the competition) to process financial PDFs, extract content from embedded images and documents, and deploy LLMs to identify specific data points like quarterly revenue.
- The end-to-end pipeline demonstrates how Daft bridges multimodal inputs and structured outputs, delivering analysis-ready data in modern table formats like Apache Iceberg—all within a single, cohesive framework.
Native model operators
Embeddings, LLM extraction, and structured outputs as first-class operations. Plug in models from OpenAI, Hugging Face, or your own.
Multimodal column types
Images, video, audio, text, and embeddings as native column types. Decode, transform, and filter them like any other column.
Local to production consistency
Define pipelines once. Run them on your laptop or scale across a cluster. Same code, no rewrites.
Managed UDF runtime
Automatic batching, retries, and error handling for model UDFs. Zero-copy execution powered by Apache Arrow.
Lower memory footprint
Run the same queries with 5x less memory than alternatives. Jobs that would OOM on Spark or Pandas just work.
Built in Rust for speed
Daft's core is written in Rust. Decode video, run transforms, and join multimodal data at TB scale without paying Python overhead.
