From raw multimodal data to training-ready datasets. In one pipeline, at any scale.

From raw multimodal data to training-ready datasets. In one pipeline, at any scale.
Process video, images, audio, and sensor data alongside structured metadata in a single dataframe.
Run GPU inference/embeddings alongside CPU decode and filter in one pipeline. Daft handles the scheduling and batching, no glue code required.
Same operations you use in Pandas or Spark: filter, transform, aggregate, write. No new framework to learn.
Process video, images, audio, and sensor data alongside structured metadata in a single dataframe.
Run GPU inference/embeddings alongside CPU decode and filter in one pipeline. Daft handles the scheduling and batching, no glue code required.
Same operations you use in Pandas or Spark: filter, transform, aggregate, write. No new framework to learn.
Use Cases
import daft
from daft.functions import file, prompt, embed_text
from pydantic import BaseModel
class Classifer(BaseModel):
title: str
summary: str
keywords: list[str]
df = (
daft.from_glob_path("s3://org-data/docs/*.pdf")
.with_column("metadata",
prompt(
messages=file(daft.col("path")),
return_format=Classifer,
model="gpt-5-mini",
provider="openai",
),
)
.with_column("abstract_embedding",
embed_text(
daft.col("metadata")["summary"],
model="text-embedding-3-large"
provider="openai",
)
)
)
# Write to Turbopuffer
df.write_turbopuffer(
namespace="org-data",
distance_metric="cosine_distance",
region="us-west-2",
)Use Cases
import daft
from daft.functions import prompt
from typing import Literal
from pydantic import BaseModel, Field
from pydantic.extra_types import phone_number, email_address
EMAIL_REGEX=r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b"
PHONE_REGEX=r"\b\d{3}-\d{3}-\d{4}\b"
# Given a sample product listing
data = daft.from_pylist([{
"id": "1234",
"listing":
"""
Vintage wooden chair $50.
Email john@email.com
Phone 123-456-7890
""",
}])
# And Desired Metadata Schema
class ProductMetadata(BaseModel):
listing: str
category: Literal["furniture", "electronics"]
price: float = Field(ge=0)
has_pii: list[Literal["email", "phone", "none"]]
# Run the pipelines
df = (
data
# Extract Metadata
.with_column("product_meta",
prompt(
daft.col("listing"),
return_format=ProductMetadata,
model="gpt-5-mini",
provider="openai",
),
)
# Redact PII
.with_column("listing_redacted",
when(
lit("email").is_in(
col("product_meta")["has_pii"]
)
).then(
col("listing").regexp_replace(
EMAIL_REGEX,
"[REDACTED_EMAIL]"
)
).when(
lit("phone").is_in(
col("product_meta")["has_pii"]
)
).then(
col("listing").regexp_replace(
PHONE_REGEX,
"[REDACTED_PHONE]"
)
).otherwise(col("listing"))
)
)
# Save results to S3
df.write_parquet("s3://org-data/safe-listings")
Use Cases
@daft.func
def extract_content(document) -> str:
...
@daft.func
def openai_o3_mini(text, prompt) -> str:
...
df = daft.from_glob_paths("s3://financials/10-k/*.pdf")
df = df.with_column("document", df["url"].download())
df = df.with_column("extracted", extract_content(df["document"]))
df = df.with_column("revenue", openai_o3_mini(df["extracted"], prompt="What is the quarterly revenue?"))
df.write_iceberg(iceberg_table)Embeddings, LLM extraction, and structured outputs as first-class operations. Plug in models from OpenAI, Hugging Face, or your own.
Images, video, audio, text, and embeddings as native column types. Decode, transform, and filter them like any other column.
Define pipelines once. Run them on your laptop or scale across a cluster. Same code, no rewrites.
Automatic batching, retries, and error handling for model UDFs. Zero-copy execution powered by Apache Arrow.
Run the same queries with 5x less memory than alternatives. Jobs that would OOM on Spark or Pandas just work.
Daft's core is written in Rust. Decode video, run transforms, and join multimodal data at TB scale without paying Python overhead.
“Daft was incredible at large volumes of abnormally shaped workloads - I pointed it at 16,000 small Parquet files in a self-hosted S3 service and it just worked! It's the data engine built for the cloud and AI workloads.”