Git 3.0's upcoming SHA-256 default will be a costly mistake

23 points by mort 5 hours ago on lobsters | 18 comments

pyfisch | 3 hours ago

The article is wrong unless I am mistaken or the git documentation is outdated. Local repositories with SHA-256 hashes will be able to push and pull from remote repositories that use SHA-1 (and vice-versa) Both hashes for each object are stored on disk and translated in the fly. This sidesteps the incompatible version issue mostly.

https://git-scm.com/docs/hash-function-transition

[OP] mort | an hour ago

It seems like there is support for a translation table indeed:

The index files support a bidirectional mapping between SHA-1 names and SHA-256 names. The lookup proceeds similarly to ordinary object lookups. For example, to convert a SHA-1 name to a SHA-256 name:

  1. Look for the object in idx files. If a match is present in the idx’s sorted list of truncated SHA-1 names, then:
    1. Read the corresponding entry in the SHA-1 name order to pack name order mapping.
    2. Read the corresponding entry in the full SHA-1 name table to verify we found the right object. If it is, then
    3. Read the corresponding entry in the full SHA-256 name table. That is the object’s SHA-256 name.

However I wonder if this is what everyone ends up doing in practice? There's little guidance on transitioning SHA-1 repos to SHA-256 repos, but the little I find (both in blog posts and when asking e.g chat bots, which I imagine a lot of people will do) results in repositories without these translation tables.

The frequently-recommended way to transition a repo from SHA-1 to SHA-256 seems to be to use git fast-export to create a repo dump, then create a new repo with extensions.objectFormat set to sha256, then use git fast-import to import. If you do this, you'll end up with a repo without the translation table. To get a repo with a translation table, you apparently need to set extensions.compatObjectFormat to sha1 before git fast-import.

I tried to experiment with this, but on my macOS/Homebrew laptop, my Fedora desktop and on an Ubuntu 26.10 container, setting extensions.compatObjectFormat to anything results in errors that "compatibility hash algorithm support requires Rust". So I can't really check how it works in practice.

In any case, I believe that unless migrations to SHA-256 always ends up with a SHA-1 translation table, enough people will create repos without the translation table by accident that this will become a huge problem in practice. You just need one dependency to mess it up for builds to start breaking after all.

From discussions I've had, this isn't something that is actually implemented. I also think it's enough of a mess that it should remain unimplemented.

"Clone the repo again" seems like a completely viable option for when you want to transition.

I should probably get around to implementing this for plan 9 git. Without the dual format support, it's probably a half day of work. With it, it's a rather complicated problem to solve.

[OP] mort | 5 hours ago

I don't know that ai agree with the author, that it's a mistake but I found the topic extremely interesting. I use a lot of software which pretty much assumes that a SHA1 is a permanent identifier of a commit; Yocto recipes, Nix, git submodules, tools like repo and gclient, etc. If repositories start rewriting their history to migrate to SHA256 and the old commits get garbage collected, this will result in massive breakage across all sorts of things and it'll be my job to find workarounds for a lot of it.

So I'm interested in seeing a lobste.rs discussion on the topic.

(Oh and it's finally an opportunity to use the merkle-trees tag! Git and other DVCSes are the only interesting applications of merkle trees as far as I'm concerned)

DVCSes are the only interesting applications of merkle trees

It's used in BitTorrent too. Not a DVCS.

[OP] mort | 21 minutes ago

Oh, thanks for the correction.

Though for whatever reason, the merkle-trees tag seems to have been removed, even though this article is entirely about merkle trees.

The tag used to have the label "blockchain", and maybe still has in the eyes of the mods.

I totally agree with @mk12 in that linked thread, it's silly to not call things by their proper name. You're right that it's appropriate on this post, and it is silly of the moderators to misuse the English language.

hailey | an hour ago

There’s one glaring hole in Scott’s “trust is in the distribution” argument specifically as it relates to GitHub which is that GitHub commingles objects across the whole fork network. Only the refs namespace is separated.

You don’t need to breach GitHub’s authentication to get a forged object into rust-lang/rust. It is much simpler - you fork the repo and push the forged object to your fork. It is then visible in the parent.

I am actually quite surprised that the article overlooks this angle given that Scott was a cofounder of GitHub.

[OP] mort | an hour ago

Yeah, and it also overlooks mirrors: one of the great things about git is that I can download a repository from any mirror and re-hash, and as long as the computed hash of a given commit is what I expect, I know that the repo I got from the mirror is good. Especially in systems like Yocto, where there are hundreds of random projects downloaded from all sorts of git repositories from lots of git hosts, it's good to know that if one project's git host gets compromised, we'll detect it because the checksums won't match up anymore.

if this vulnerability is already exploitable today, why is it not actively exploited this way?

[OP] mort | 22 minutes ago

SHA-1 is semi-broken. There have been collisions found and there have been found cryptographic weaknesses which makes it significantly weaker than the number of bits would suggest. But it's still not by any means easy to forge a commit which both has the malicious payload you want and has a hash collision with a relevant commit. The collisions which have been made so far have very much been proof-of-concept.

But, like with post-quantum encryption algorithms, it's probably better to transition before it becomes a big enough problem that there are exploits in the wild.

Also, Git and Git hosts today use what they call "sha1dc"; sha1 with collision detection. I don't understand how the cryptography works but it apparently is able to somewhat work around the specific issues which have been used to produce collisions so far. So pushing e.g two colliding PDFs to GitHub just won't work these days.

I had this same experience working on a proprietary backup system that used hashes heavily in 2016 or so. “FIPS says we can only use SHA-256.” “But we use hashes for deduplication, not security.” “Doesn’t matter, FIPS says it.” “But we have thousands of users. Many are on older CPUs where that hash is expensive.” “Doesn’t matter.” “Can we just throw truncated SHA-256 in the same field as SHA-1 and call it a day?” “No.” Anyway we upgraded all our users to SHA-256.

elobdog | 2 hours ago

That's a lost battle. The "Compliance" world freaks out when they see SHA-1, even when you do not use it for any cryptographically relevant purpose (like signatures).

quasi_qua_quasi | an hour ago

I had a very funny inverse story of this recently, where the "security" training at work suggested I verify the authenticity of files using MD5...

In the case of Git though, the existence and implementation of commit signing does require security properties of the object and commit hashes.

I think that reasoning led to a svn repository being catastrophically corrupted by the shattered sha-1 collision.

It’s a common fallacy that deduplication doesn’t depend on the security properties of cryptographic hash functions: if the system doesn’t have any extra machinery to handle a collision then a broken hash function will cause loss of data integrity and probably availability, two of the three legs of the CIA triad.

SamRW | an hour ago

I do appreciate making this change will likely be painful, but I would hesitate to downplay the impact of these collisions. In these cases, I'd value the opinions of cryptographers, who have been pointing the various deficiencies in SHA1 for an extended period.

interlandi | 18 minutes ago

I am no cryptographer, nor an expert on the "public health" side of software engineering, but I would wager that this decision is supposed to be forward-looking rather that trying to mitigate some non-existent immediate issue. I do not like that the article positions the change as the latter (to an extent, I know they offered more than that). 15 years from now, some number of the world's repositories will still be using SHA-1. Nobody can say for certainty that SHA-1 will not be broken in 15 years. That number will be smaller then, given that they are making the switch now, than if Git had waited. One thing I would also add is that adding additional hash functions on top of existing ones may not be suitable solution, since, at least intuitively, it must increase both the cryptographic integrity and the vulnerable surface simultaneously.