SAMSKARA

One workflow, not a feature list.

Every SAMSKARA capability exists to move data through the same six stages — the stages your team already works in.

Connect→Explore→Transform→Schedule→Monitor→Share
Connect

Bring your own storage, databases, and runtime.

Point SAMSKARA at Postgres, MySQL, object storage, or your existing warehouse via runtime, storage, and database profiles. SAMSKARA is the control plane — it doesn't lock your data into a walled garden.

SAMSKARA AnyBase connection list showing MySQL and ArrowLake connections
Explore

One catalog tree for everything.

The Unified Data Explorer puts ArrowLake tables and external sources side by side — schema, snapshots, AI-generated insights, and column-level tags/comments in one place.

Unified Data Explorer showing the ArrowLake catalog tree and a table's schema with generated read/write snippets
Transform

Notebooks and SQL, not YAML.

Interactive Python notebooks with per-cell run/stop and session isolation, plus a Monaco-powered SQL Workbench with AI-assisted text-to-SQL. Your team writes the logic they already know.

SAMSKARA notebook with executed cells and the AI Sparkle panel open
Schedule

Cron-based orchestration that just works.

Dependent job chaining, live countdowns, run history with full cell code and output, webhook triggers, and per-org concurrency limits — no separate orchestrator to stand up.

SAMSKARA Scheduler showing active scheduled notebook jobs with cron schedules and run history
Monitor

Dashboards your team builds themselves.

SQL-driven charts — line, bar, area, pie, metric, table — with variables and per-widget auto-refresh. No separate BI tool procurement.

A SAMSKARA analytics dashboard with line, bar, pie, and table widgets built from a weather dataset
Share

Governance and access control, built in.

Org roles, group-based catalog grants, dashboard sharing, and audit trails come standard — so opening SAMSKARA up to more of the team doesn't mean giving up control.

SAMSKARA Access Control settings showing user groups for bulk permission management

More built-in capabilities

Everything ships in the box — no plugin marketplace, no add-ons.

Show allHide
  • ⚡

    Parallel execution

    sm.parallel_map fans out any function across a list of inputs with automatic thread scaling and per-item error isolation.

  • 📄

    PDF data extraction

    sm.extract_pdf pulls structured data from invoices, COA documents, and reports straight into a DataFrame.

  • 📂

    File ingest & Auto Ingest

    Upload CSV, Excel, Parquet, JSON, or XML and land it as an Iceberg table in one step, or schedule a recurring rule to watch a storage folder and ingest new files automatically — permissions re-checked on every run.

  • 📡

    AnyStream

    Read and write Kafka, Redpanda, Pulsar, NATS, or Kinesis topics directly from a notebook, or run continuous background ingestion into ArrowLake — no separate stream-processing cluster.

  • 🚀

    Environment Promotion

    Promote a notebook from Development to Production with one click — its code travels, while database connections, storage profiles, and secrets re-resolve against the target environment by name.

  • 🔍

    AI Insights & Lineage

    The Unified Data Explorer surfaces AI-generated table summaries, anomaly flags, and a full read/write history per table.

  • 🌐

    Data acquisition

    sm.fetch() and sm.extract() handle retries, pagination, encoding detection, and API auth — no raw HTTP library boilerplate.

  • 📊

    BI connectivity

    Power BI, Tableau, and Excel query ArrowLake via a read-only REST SQL API authenticated with personal access tokens.

  • 🔀

    Git integration & CI/CD

    Push and pull notebooks to any Git repository with per-cell diff review, or wire a CI/CD pipeline to promote commits straight into production via a scoped, revocable service token.

  • 🔐

    Encrypted secrets

    Org-scoped key-value store; secrets are injected into notebook runs as sm.secrets["KEY"] — never stored in plain text.

  • 🏔️

    Apache Iceberg / Polaris

    ArrowLake stores data in open Iceberg format. Migrate to Apache Polaris for enterprise catalog governance without moving any data.

  • ⏱️

    Time-travel queries

    Read any ArrowLake table as of a previous snapshot or timestamp — roll back to yesterday's state with one parameter.

  • 🔁

    Upsert / merge

    sm.merge_arrowdelta handles SCD Type 1 and Type 2 merge-by-key patterns on top of Iceberg merge-on-read.

  • 🏷️

    Column-level tags & comments

    Annotate tables and columns in the Data Explorer; metadata is stored as Iceberg table properties and survives catalog operations.

  • 🐳

    Docker runtime isolation

    Each notebook session runs in its own container — no cross-session package conflicts, no dependency drift between projects.

  • 👥

    Group-based access control

    Assign users to groups; grant catalog, dashboard, and scheduler access to groups instead of individuals.

  • 📬

    Invite-only access

    Disable public registration and route new users through an admin-approved access request flow with automatic email notification.

Ready to see your own workflow in SAMSKARA?