Self-hosted data engineering
Modern data engineering without platform complexity.
Samskara is the self-hosted data engineering platform for teams that have outgrown scripts but don't need the complexity and cost of a large cloud data platform.

One workflow, start to finish
Everything your data team needs, without the platform tax
No Spark clusters to babysit, no six-figure platform contract. Just the tools a 20–200 person team actually uses, running on infrastructure you already control.
Notebook Execution
Python notebooks with per-cell run/stop, isolated sessions, and full artifact traces.
SQL Workbench
Monaco-powered SQL editor with schema browsing and AI-assisted text-to-SQL.
ArrowLake
DuckDB + Iceberg + object storage, with time travel and no separate lakehouse cluster to run.
Scheduler
Cron-based orchestration with dependent job chaining and live run history.
Dashboards
SQL-driven charts and metrics your team can build without a BI team.
AI Sparkle
Notebook- and schema-aware AI that learns from your team's own successful runs.
Governance & RBAC
Org roles, catalog-level grants, audit trails — built in, not bolted on.
Git Integration
Version-controlled notebooks and projects, the way engineers already work.
Real projects, not slideware
Every capability on this site has been proven end-to-end on real data projects built in SAMSKARA.
Weather
Ingest, transform, and dashboard public weather data end-to-end.
Public Health
Multi-source health datasets landed in ArrowLake and scheduled nightly.
Northwind
The classic sample database, rebuilt as a real SAMSKARA project.
See it running on your data
30 minutes, your use case, no slideware.