Skip to main content

Batch APIs

The Python binding provides nine batch APIs that fan out multiple RPCs with bounded concurrency (at most MAX_BATCH_RPC_IN_FLIGHT = 64 in flight). Each batch completes in a single PyO3 boundary crossing, eliminating per-call GIL acquisition and making them dramatically faster than N individual calls under GIL contention.

When to Use Batch APIs​

ScenarioUse batch?Why
Check existence of 100+ pathsbatch_existsOne GIL crossing instead of 100
Fetch metadata for a file listbatch_get_statusResults in input order, bounded fan-out
Create / delete / rename a batchbatch_create_* / batch_delete / batch_renameConcurrent ops; partial changes may remain on failure
Open 50 files for parallel readbatch_open_fileStreams cleaned up on partial failure
List multiple directoriesbatch_list_status / batch_list_status_groupedConcurrent listing
warning

All batch APIs fail the whole batch on the first error, but operations that already completed are not rolled back. Use individual calls if you need per-path error isolation or atomicity.

Status APIs​

batch_exists(paths) → list[bool]​

paths = ["/data/a", "/data/b", "/data/missing"]
results = await fs.batch_exists(paths)
# [True, True, False] — in input order

batch_get_status(paths) → list[URIStatus]​

statuses = await fs.batch_get_status(["/data/a", "/data/b"])
for s in statuses:
print(f"{s.path}: {s.length} bytes")

Raises NotFound if any path is missing — the whole batch fails.

list_status_grouped(path) → URIStatusList (lazy)​

grouped = await fs.list_status_grouped("/data", recursive=False)
print(len(grouped)) # O(1) — no URIStatus objects created
first = grouped[0] # materialises one URIStatus on demand
for entry in grouped: # iteration materialises one at a time
print(entry.name)

URIStatusList is a lazy container: len() is O(1) with zero object creation, and __getitem__ / __iter__ materialise URIStatus objects on demand. This reduces GIL occupancy by ~99% for N=100 entries.

batch_list_status_grouped(dirs) → list[URIStatusList]​

dirs = ["/data/d1", "/data/d2", "/data/d3"]
groups = await fs.batch_list_status_grouped(dirs, recursive=False)
for i, g in enumerate(groups):
print(f"dir {i}: {len(g)} entries")

Lazy counterpart to batch_list_status. Each directory's entries are returned as a URIStatusList (1 Python object per directory) instead of list[URIStatus] (N objects per directory).

batch_list_status(dirs) → list[list[URIStatus]]​

Eager variant — materialises all entries immediately. Use when you need a plain list[URIStatus] for slicing or library interop.

File-Lifecycle APIs​

batch_create_file(paths) → list[int]​

Creates and closes an empty file at every path. Returns bytes written per file (always 0 for empty files) in input order.

files = ["/data/f1", "/data/f2", "/data/f3"]
written = await fs.batch_create_file(files)
# [0, 0, 0]

batch_create_dir(paths) → None​

dirs = ["/data/d1", "/data/d2"]
await fs.batch_create_dir(dirs, recursive=True)

batch_rename(pairs) → None​

pairs is a flat list of alternating source and destination: [src_0, dst_0, src_1, dst_1, ...]. Length must be even.

await fs.batch_rename(["/data/old1", "/data/new1", "/data/old2", "/data/new2"])

Raises ValueError if the list length is odd.

batch_delete(paths) → None​

await fs.batch_delete(["/data/d1", "/data/d2"], recursive=True)

Options: recursive, unchecked (skip empty-check), goosefs_only (don't propagate to UFS).

tip

When batch-deleting a tree, only include the parent directories (with recursive=True). Including child files in the same batch can race: if a parent finishes first, the child delete sees a missing path and fails the whole batch.

batch_open_file(paths) → list[AsyncFileReader]​

Opens N read streams concurrently. Returns readers in input order.

readers = await fs.batch_open_file(["/data/a", "/data/b", "/data/c"])
for r in readers:
data = await r.read()
print(len(data))

Resource cleanup on partial failure: after all dispatched open attempts complete, if any path failed, all successfully opened streams are dropped (their Drop releases worker resources) and the batch raises an error.

Sync API​

Every batch API is available on the synchronous Goosefs wrapper, with the same names and argument shapes:

with Goosefs(cfg) as fs:
results = fs.batch_exists(["/data/a", "/data/b"])
fs.batch_create_dir(["/data/d1", "/data/d2"])

# Returns synchronous FileReader objects — no asyncio runtime needed.
readers = fs.batch_open_file(["/data/a", "/data/b"])
try:
for r in readers:
print(len(r.read()))
finally:
for r in readers:
r.close()

Performance Characteristics​

Batch APIs complete in a single PyO3 boundary crossing instead of N. For a directory with N=100 entries, list_status_grouped creates 1 Python object (~0.3 µs GIL occupancy) instead of 100 URIStatus objects (~33.4 µs total), reducing GIL occupancy by ~99%. The speedup comes from eliminating per-call PyO3 boundary crossings (each briefly acquires the GIL and allocates Python objects).