One workflow, not a feature list.
Every SAMSKARA capability exists to move data through the same six stages — the stages your team already works in.
Bring your own storage, databases, and runtime.
Point SAMSKARA at Postgres, MySQL, object storage, or your existing warehouse via runtime, storage, and database profiles. SAMSKARA is the control plane — it doesn't lock your data into a walled garden.

One catalog tree for everything.
The Unified Data Explorer puts ArrowLake tables and external sources side by side — schema, snapshots, AI-generated insights, and column-level tags/comments in one place.

Notebooks and SQL, not YAML.
Interactive Python notebooks with per-cell run/stop and session isolation, plus a Monaco-powered SQL Workbench with AI-assisted text-to-SQL. Your team writes the logic they already know.

Cron-based orchestration that just works.
Dependent job chaining, live countdowns, run history with full cell code and output, webhook triggers, and per-org concurrency limits — no separate orchestrator to stand up.

Dashboards your team builds themselves.
SQL-driven charts — line, bar, area, pie, metric, table — with variables and per-widget auto-refresh. No separate BI tool procurement.

Governance and access control, built in.
Org roles, group-based catalog grants, dashboard sharing, and audit trails come standard — so opening SAMSKARA up to more of the team doesn't mean giving up control.

More built-in capabilities
Everything ships in the box — no plugin marketplace, no add-ons.
Show allHide
More built-in capabilities
Everything ships in the box — no plugin marketplace, no add-ons.
- ⚡
Parallel execution
sm.parallel_map fans out any function across a list of inputs with automatic thread scaling and per-item error isolation.
- 📄
PDF data extraction
sm.extract_pdf pulls structured data from invoices, COA documents, and reports straight into a DataFrame.
- 📂
File ingest & Auto Ingest
Upload CSV, Excel, Parquet, JSON, or XML and land it as an Iceberg table in one step, or schedule a recurring rule to watch a storage folder and ingest new files automatically — permissions re-checked on every run.
- 📡
AnyStream
Read and write Kafka, Redpanda, Pulsar, NATS, or Kinesis topics directly from a notebook, or run continuous background ingestion into ArrowLake — no separate stream-processing cluster.
- 🚀
Environment Promotion
Promote a notebook from Development to Production with one click — its code travels, while database connections, storage profiles, and secrets re-resolve against the target environment by name.
- 🔍
AI Insights & Lineage
The Unified Data Explorer surfaces AI-generated table summaries, anomaly flags, and a full read/write history per table.
- 🌐
Data acquisition
sm.fetch() and sm.extract() handle retries, pagination, encoding detection, and API auth — no raw HTTP library boilerplate.
- 📊
BI connectivity
Power BI, Tableau, and Excel query ArrowLake via a read-only REST SQL API authenticated with personal access tokens.
- 🔀
Git integration & CI/CD
Push and pull notebooks to any Git repository with per-cell diff review, or wire a CI/CD pipeline to promote commits straight into production via a scoped, revocable service token.
- 🔐
Encrypted secrets
Org-scoped key-value store; secrets are injected into notebook runs as sm.secrets["KEY"] — never stored in plain text.
- 🏔️
Apache Iceberg / Polaris
ArrowLake stores data in open Iceberg format. Migrate to Apache Polaris for enterprise catalog governance without moving any data.
- ⏱️
Time-travel queries
Read any ArrowLake table as of a previous snapshot or timestamp — roll back to yesterday's state with one parameter.
- 🔁
Upsert / merge
sm.merge_arrowdelta handles SCD Type 1 and Type 2 merge-by-key patterns on top of Iceberg merge-on-read.
- 🏷️
Column-level tags & comments
Annotate tables and columns in the Data Explorer; metadata is stored as Iceberg table properties and survives catalog operations.
- 🐳
Docker runtime isolation
Each notebook session runs in its own container — no cross-session package conflicts, no dependency drift between projects.
- 👥
Group-based access control
Assign users to groups; grant catalog, dashboard, and scheduler access to groups instead of individuals.
- 📬
Invite-only access
Disable public registration and route new users through an admin-approved access request flow with automatic email notification.