The hidden design compromises of Docker layers

5 points by loige 13 hours ago on lobsters | 10 comments

mcherm | 8 hours ago

Wow -- that was a VERY long writeup. I feel like it could have been covered in 2 paragraphs at most. (It WAS an interesting implementation detail of Docker layers.)

[OP] loige | 7 hours ago

ahah fair enough! I should have probably had a TLDR; at the top and then a disclaimer "keep reading ONLY if you are really really bored!"

pronoiac | 5 hours ago

The Cheat Sheet section, with a bit of context, could be that. I think it left out "white out this directory."

I'm still curious about seekable tar files and lazy-loading.

mcherm | 5 hours ago

Honestly, that would have been hilarious! 😆

david_chisnall | 6 hours ago

Thanks @loige, this is a great article (well written and informative)!

One thing to mention from the terminology:

Whiteout is a really common term for any kind of merged dictionary (filesystem or otherwise) where a newer layer needs a mechanism for deleting an entry from an older one. And it's another one of these culture-specific terms leaking into generic discourse. Wite-Out is a brand of correctional fluid sold in the USA. For folks in Europe, Tipp-Ex is the equivalent.

For those people too young to have had to suffer with typewriters or pens and who have therefore never had to deal with either of these products:

When you wrote with a pen or typed with a typewriter, there was no delete functionality. To do the equivalent of a deletion, you painted some white fluid over the top (the typewriter version involved little white ink pads that fitted in front of the ribbon so you could type a white letter over the black one and 'delete' it). This was picked up as a metaphor in computing. The US-trademarked term was changed to its more conventional spelling.

Naming is hard.

Locally, the storage driver (for example, overlay2) keeps each layer as an extracted directory and uses a union filesystem to stack them

It's worth noting that this, as an abstraction, was one of the big changes with containerd and OCI standardisation for this model. Docker's filesystem abstraction explicitly did work like (I think it can now do both?) this but later implementations converged on a snapshotter abstraction rather than an overlay abstraction.

In the overlay model, you create a bunch of directories and then assemble them into layers. In the snapshotter model, you create a directory tree then apply a bunch of changes and assign that a name, then you apply other changes and assign it a different name. The overlay model is more flexible: It can express everything that you can express in the snapshotter model and also some other things such as different composition orders (so, in theory, could provide mixins rather than strict sequential deltas). The problem is that a bunch of the desirable implementation substrates are snapshots, not arbitrary compositions. Implementing the simpler mode with the more complex is easy. It's easy to use an overlay filesystem to implement a snapshot model, it's hard to use a CoW filesystem to implement the overlay model.

This is how a modification is represented: the new layer simply ships a new, complete version of the file

This is one of the places where the article's earlier comment about the difference between the snapshotter and the distribution format being different is important. If you use an overlay FS driver, you'll have two directories with different versions of the file. If you use the ZFS snapshotter, you'll extract the second one on top of the first and, if the changes are small, you may or may not get a separate copy depending on whether you've enabled dedup on the datasets that it uses.

BTW, am I the weird one, or did your brain come up with the same question? 🧠

Both.

In other words, PAX is tar’s official escape hatch for extra metadata

Pax isn't an acronym, it literally means peace (in Latin) because it's the result of a compromise between the divergence in the UNIX world between tar and cpio (and their various forks). Ironically, POSIX 2001 defined the modern pax format, which is supported by tar on GNU and FreeBSD systems but not by pax. Because consensus is even harder than naming.

As files prefixed with .wh. are special whiteout markers, it is not possible to create a filesystem which has a file or directory with a name beginning with .wh..

This seems weird, but there are a bunch of other filesystem-specific limitations that OCI container authors need to care about, so it's probably not the weirdest. Try creating a Windows OCI container image containing a file called COM1 some time.

icefox | 7 hours ago

Hm, good writeup. I feel like "diff of a filesystem" is a wheel that's been reinvented enough times we should have an off the shelf solution for it by now...

[OP] loige | 7 hours ago

Indeed, I was more interested in the overall design choices for this one. Looks like they went for the simplest yet effective enough combo!

k749gtnc9l3w | 7 hours ago

By the way, deletion-marking files with a special naming convention are also used in various versions of UnionFS, since before Docker.

[OP] loige | 5 hours ago

Indeed and thanks for mentioning it! I think that's an interesting bit of the history which i am pretty sure inspired the current design of docker layers... I did find out about this during my research and I am pretty sure i mentioned it somewhere in the article, since I kept wondering "why didn't they just use Tar PAX headers??"

k749gtnc9l3w | 3 hours ago

What do you mean inspired, aren't your currently included quotes imply it was forced to go this way if it wanted to use (to unpack the base layer once and use it in many different containers) the only in-kernel union FS available back then?