File scanner for multi-part splitting
Core Engineering (Flexgroups), ONTAP filesystem — Linux
Implemented the scanner that decides which files are worth splitting into parts, and tuned its default behaviour for scan time and efficiency across the installed fleet.
The multi-part format is not free. Splitting a file into parts buys concurrency and large-file support, but it costs metadata, adds indirection, and makes every access a little more work. Applying it to every file would be a bad trade.
So the system needs a scanner: something that looks at a filesystem and decides which files would actually benefit from being split.
The scanner
I implemented the scanner that searches for potential files that would benefit the user most when splitting, and I tuned the default scanner for time and efficiency.
The phrase "benefit the user most" is doing real work in that sentence. The scanner is not a correctness component — the filesystem is correct either way. It is a judgement call about where to spend overhead, and getting it wrong in either direction has a real cost:
- Too eager, and you pay the multi-part cost on files that gained nothing. The system does more work for no user-visible benefit.
- Too conservative, and the files most in need of splitting — the large, concurrently accessed ones — keep suffering, and the feature is invisible to exactly the users it was built for.
Why scan time is the thing to tune
The obvious way to scan is to look at everything. The problem is the shape of the deployment: this runs across an installed fleet, so the scan is not a one-time offline job. It runs while the system is live, on filesystems that are actively changing underneath it.
That constrains the design in ways that are easy to underestimate:
- A scan that takes too long is functionally a scan that does not run. Past some duration it stops being worth the I/O, and the tuning target is the point where the remaining candidates stop justifying the cost.
- The filesystem is moving while you scan. Entries appear, change, and disappear. Handling that without producing a wrong result is harder than the scanning itself, and "we got it slightly wrong" means splitting a file mid-write.
- You have to be able to explain the result. A split decision is a durable change to a user's data. Whatever the heuristic picks, there needs to be a reason attached to it, because "the scanner decided" is not an answer anyone can act on when a file splits badly.
So tuning the default was an exercise in finding the point where the scan is cheap enough to run often, thorough enough to catch the files that matter, and conservative enough that a bad call is rare and reversible.
What it taught me
This is the work that most shaped how I approach performance problems. The temptation in a filesystem is to optimize the thing you can measure — bytes per second, cache hit rate — and assume the result is better. Here the measurable quantity and the goal had come apart: a faster scan is not the objective, a scan that runs often enough to be useful while finding files worth acting on is. Tuning toward the first one would have made the second one worse.