Why Samskara?
Samskara is the self-hosted data engineering platform for teams that have outgrown scripts but don't need the complexity and cost of a large cloud data platform. Here's what that means in practice.
“I don't want to maintain Spark clusters.”
SAMSKARA runs on DuckDB and Iceberg over object storage. No cluster to size, patch, or babysit — ArrowLake gives you Iceberg time travel and SQL catalog semantics with no separate lakehouse infrastructure to operate.
“I don't need an enterprise cloud data platform for a 20 person company.”
Most teams this size are paying for elasticity they'll never use. SAMSKARA is self-hosted and sized for the workloads a small data team actually runs — not a platform built for a thousand-person org.
“My team just wants notebooks and SQL.”
That's the whole surface area: a Python notebook IDE and a SQL Workbench, both wired directly into the same catalog, scheduler, and dashboards. No new abstraction layer to learn.
“I need governance.”
Org roles, group-based catalog grants, project-scoped secrets, and audit trails are built into the core — not a separate enterprise SKU you have to negotiate for.
“I don't want another vendor holding my data hostage.”
SAMSKARA is the control plane, not the storage. Your data lives in your object storage and your databases; SAMSKARA orchestrates and executes against them.
“I need this running in an afternoon, not a quarter.”
Docker Compose brings up Postgres, Redis, the API, worker, and frontend together. There's no data platform migration project required to get started.
Sound familiar?
Let's talk about your team's data stack.