SAMSKARA

September 27, 2026

Why we built ArrowLake's maintenance lifecycle

When we wrote about why we built ArrowLake, we talked about what we gave up by not running a distributed engine or a dedicated metastore cluster. What we didn’t talk about yet is what that trade-off costs you later: a managed lakehouse has a vendor quietly running compaction and garbage collection behind the scenes. A self-managed one doesn’t, unless you build that yourself. Over the last several weeks, we did.

The shape of the problem

Every append to an Iceberg table adds a manifest. Enough small appends and a table accumulates hundreds of manifests before it accumulates much data — reads get slower reading metadata, not rows. Enough small files and the same problem shows up on the data side. Old snapshots pile up too: useful for time travel, mostly useless after that, and each one keeps its files alive in the table’s history even after nobody could reasonably want them back.

None of this is exotic — it’s the shape every Iceberg deployment eventually hits. What’s different building it inside SAMSKARA is that table health lives right where you already look at your tables, backed by a real audit trail and locking discipline, instead of a cron job someone wrote once and nobody remembers restarting.

Health first, action second

A table’s Health tab reports the signals that matter — manifest count, small-file ratio, snapshot count — and turns them into concrete recommendations: rewrite manifests, compact data files, expire old snapshots. Each became its own operation, individually triggerable and individually measurable, with a before/after health snapshot recorded every time it runs. A table under maintenance gets locked for the duration, so two operations can’t collide, and a long-running compaction hands back a job id immediately instead of holding an HTTP request open until it finishes.

Once that existed by hand, the obvious next step was letting it run itself: an ArrowLake profile can opt into scheduled evaluation, gated by a per-org kill switch that decides whether the scheduler is allowed to actually execute anything or only recommend it.

The operation we didn’t ship casually

Snapshot expiration only removes metadata — it drops old snapshot entries from a table’s history, but Iceberg doesn’t delete the data and manifest files those snapshots pointed to on its own. Storage never actually shrinks. Closing that gap meant writing the one ArrowLake operation capable of something every operation before it couldn’t: permanently, irreversibly deleting real files.

We didn’t treat that like just another maintenance operation. Detecting an orphan means walking every snapshot a table still has, not just the current one, and diffing that against what’s physically sitting in storage — a file is only a candidate if nothing, including an old branch or tag, still points to it. Deleting one requires organization-admin specifically, on top of the permission that gates every other maintenance action, gated again by its own kill switch that’s off by default for every organization, and re-checked live at the moment deletion actually runs rather than only when it was requested, because either can change in the gap between the two. Even a file a human has already reviewed and approved gets re-verified against a fresh scan immediately before it’s removed, so nothing is deleted that stopped being an orphan in the meantime.

Before any of that touched real data, we proved it end-to-end against a disposable table built to create a genuine orphan and then verify every safety property against it: still-referenced files correctly protected, expired files correctly detected, a deliberately stale confirmation correctly excluded, and a live read afterward returning exactly the rows that were supposed to survive.

Where it’s headed

A UI for this — reviewing candidates and flipping the kill switch without a script — is next. So is a related gap we found along the way: dropping a table today only removes its catalog entry, not its files, so a dropped table’s storage sits there forever. Fixing that gets its own review, on top of the same safety pattern this work just proved out.

If you’re running Iceberg without a cluster and wondering what actually keeps it healthy over time, take a look at ArrowLake or sign up free and watch a table’s Health tab tell you what it needs.