{"id":67031,"date":"2022-09-13T09:00:49","date_gmt":"2022-09-13T16:00:49","guid":{"rendered":"https:\/\/github.blog\/?p=67031"},"modified":"2022-09-13T09:54:52","modified_gmt":"2022-09-13T16:54:52","slug":"scaling-gits-garbage-collection","status":"publish","type":"post","link":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/","title":{"rendered":"Scaling Git\u2019s garbage collection"},"content":{"rendered":"<p>At GitHub, we store a lot of Git data: more than 18.6 petabytes of it, to be precise. That&#8217;s more than six times the size of the Library of Congress&#8217;s digital collections<sup id=\"fnref-67031-1\"><a href=\"#fn-67031-1\" class=\"jetpack-footnote\" title=\"Read footnote.\">1<\/a><\/sup>. Most of that data comes from the contents of your repositories: your READMEs, source files, tests, licenses, and so on.<\/p>\n<p>But some of that data is just junk: some bit of your repository that is no longer important. It could be a file that you <a href=\"https:\/\/git-scm.com\/docs\/git-push#Documentation\/git-push.txt---force\">force-pushed<\/a> over, or the contents of a branch you deleted without merging. In general, this slice of repository data is anything that isn&#8217;t contained in at least one of your repository&#8217;s branches or tags. Normally, we don&#8217;t remove any unreachable data from repositories. But occasionally we do, usually <a href=\"https:\/\/docs.github.com\/en\/authentication\/keeping-your-account-and-data-secure\/removing-sensitive-data-from-a-repository#fully-removing-the-data-from-github\">to remove sensitive data, like passwords or SSH keys<\/a> from your repository&#8217;s history.<\/p>\n<p>The process for permanently removing unreachable objects from a repository&#8217;s history has a history of causing problems within GitHub, especially in busy repositories or ones with lots of objects. In this post, we&#8217;ll talk about what those problems were, why we had them, the <a href=\"https:\/\/git-scm.com\/docs\/cruft-packs\">tools we built<\/a> to address them, and some interesting ways we&#8217;ve built on top of them. All of this work was <a href=\"https:\/\/github.com\/git\/git\/compare\/37d4ae58efcc9f716435f327c39d5552aedb4b7c...a613164257b46700ca583bdcab160c712ad392fe\">contributed upstream<\/a> to the open-source Git project. Let&#8217;s dive in.<\/p>\n<h2 id=\"object-reachability\"><a class=\"heading-link\" href=\"#object-reachability\">Object reachability<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>In this post, we&#8217;re going to talk a lot about &#8220;reachable&#8221; and &#8220;unreachable&#8221; objects. You may have heard these terms before, but perhaps only casually. Since we&#8217;re going to use them a lot, it will help to have more concrete definitions of the two. An object is <em>reachable<\/em> when there is at least one branch or tag along which you can reach the object in question. An object is &#8220;reached&#8221; by crawling through history\u2014from commits to their parents, commits to their root trees, and trees to their sub-trees and blobs. An object is <em>unreachable<\/em> when no such branch or tag exists.<\/p>\n<p><img data-recalc-dims=\"1\" decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=1024&#038;resize=1024%2C493\" alt=\"Sample object graph showing commits, with arrows connecting them to their parents. A few commits have boxes that are connected to them, which represent the tips of branches and tags.\" width=\"1024\" height=\"493\" class=\"aligncenter size-large wp-image-67035 width-fit\" srcset=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=1600 1600w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=300 300w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=768 768w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=1024 1024w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection2.png?w=1536 1536w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>Here, we&#8217;re looking at a sample object graph. For simplicity, I&#8217;m only showing commits (identified here as circles). Arrows point from commits to their parent(s). A few commits have boxes that are connected to them, which represent the tips of branches and tags.<\/p>\n<p>The parts of the graph that are colored blue are reachable, and the red parts are considered unreachable. You&#8217;ll find that if you start at any branch or tag, and follow its arrows, that all commits along that path are considered reachable. Note that unreachable commits which have reachable ones as parents (in our diagram above, anytime an arrow points from a red commit to a blue one) are still considered unreachable, since they are not contained within any branch or tag.<\/p>\n<p>Unreachable objects can also appear in clusters that are totally disconnected from the main object graph, as indicated by the two lone red commits towards the right-hand side of the image.<\/p>\n<h2 id=\"pruning-unreachable-objects\"><a class=\"heading-link\" href=\"#pruning-unreachable-objects\">Pruning unreachable objects<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>Normally, unreachable objects stick around in your repository until they are either automatically or manually cleaned up. If you&#8217;ve ever seen the message, &#8220;Auto packing the repository for optimum performance,&#8221; in your terminal, Git is doing this for you in the background. You can also trigger <a href=\"https:\/\/en.wikipedia.org\/wiki\/Garbage_collection_(computer_science)\">garbage collection<\/a> manually by running:<\/p>\n<pre><code>$ git gc --prune=&lt;date&gt;\n<\/code><\/pre>\n<p>That tells Git to trigger a garbage collection and remove unreachable objects. But observant readers might notice the optional <code>&lt;date&gt;<\/code> parameter to the <code>--prune<\/code> flag. What is that? The short answer is that Git allows you to restrict which objects get permanently deleted based on the last time they were written. But to fully explain, we first need to talk a little bit about a race condition that can occur when removing objects from a Git repository.<\/p>\n<h3 id=\"object-deletion-raciness\"><a class=\"heading-link\" href=\"#object-deletion-raciness\">Object deletion raciness<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Normally, deleting an unreachable object from a Git repository should not be a notable event. Since the object is unreachable, it&#8217;s not part of any branch or tag, and so deleting it doesn&#8217;t change the repository&#8217;s reachable state. In other words, removing an unreachable object from a repository should be as simple as:<\/p>\n<ol>\n<li>Repacking the repository to remove any copies of the object in question (and recomputing any deltas that are based on that object).<\/li>\n<li>Removing any loose copies of the object that happen to exist.<\/li>\n<li>Updating any additional indexes (like the <code><a href=\"https:\/\/git-scm.com\/docs\/multi-pack-index\">multi-pack index<\/a><\/code>, or <code><a href=\"https:\/\/git-scm.com\/docs\/commit-graph\">commit-graph<\/a><\/code>) that depend on the (now stale) packs that were removed.<\/li>\n<\/ol>\n<p>The racy behavior occurs when a repository receives one or more pushes during this process. The main culprit is that the <a href=\"https:\/\/git-scm.com\/docs\/protocol-v2\">server advertises its objects<\/a> at a different point in time from processing the objects that the client sent based on that advertisement.<\/p>\n<p>Consider what happens if Git decides (as part of running a <code>git gc<\/code> operation) that it wants to delete some unreachable object <code>C<\/code>. If <code>C<\/code> becomes reachable by some background reference update (e.g., an incoming push that creates a new branch pointing at <code>C<\/code>), it will then be advertised to any incoming pushes. If one of these pushes happens before <code>C<\/code> is actually removed, then the repository can end up in a corrupt state. Since the pusher will assume <code>C<\/code> is reachable (since it was part of the object advertisement), it is allowed to include objects that either reference or depend on <code>C<\/code>, without sending <code>C<\/code> itself. If <code>C<\/code> is then deleted while other reachable parts of the repository depend on it, then the repository will be left in a corrupt state.<\/p>\n<p>Suppose the server receives that push before proceeding to delete <code>C<\/code>. Then, any objects from the incoming push that are related to it would be immediately corrupt. Reachable parts of the repository that reference <code>C<\/code> are no longer <a href=\"https:\/\/en.wikipedia.org\/wiki\/Closure_(mathematics)\">closed<\/a><sup id=\"fnref-67031-2\"><a href=\"#fn-67031-2\" class=\"jetpack-footnote\" title=\"Read footnote.\">2<\/a><\/sup> over reachability since <code>C<\/code> is missing. And any objects that are stored as a delta against <code>C<\/code> can no longer be inflated for the same reason.<\/p>\n<p><img data-recalc-dims=\"1\" decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=1024&#038;resize=1024%2C438\" alt=\"Figure demonstrating that one side (responsible for garbage collecting the repository) decides that a certain object is unreachable, while another side makes that object reachable and accepts an incoming push based on that object\u2014before the original side ultimately deletes that (now-reachable) object\u2014leaving the repository in a corrupt state.\" width=\"1024\" height=\"438\" class=\"aligncenter size-large wp-image-67036 width-fit\" srcset=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=1600 1600w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=300 300w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=768 768w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=1024 1024w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection3.png?w=1536 1536w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>In case that was confusing, the above figure should help clear things up. The general idea is that one side (responsible for garbage collecting the repository) decides that a certain object is unreachable, while another side makes that object reachable and accepts an incoming push based on that object\u2014before the original side ultimately deletes that (now-reachable) object\u2014leaving the repository in a corrupt state.<\/p>\n<h3 id=\"mitigating-object-deletion-raciness\"><a class=\"heading-link\" href=\"#mitigating-object-deletion-raciness\">Mitigating object deletion raciness<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Git does not completely prevent this race from happening. Instead, it works around the race by gradually expiring unreachable objects based on the last time they were written. This explains the mysterious <code>--prune=&lt;date&gt;<\/code> option from a few sections ago: when garbage collecting a repository, only unreachable objects which haven&#8217;t been written since <code>&lt;date&gt;<\/code> are removed. Anything else (that is, the set of objects that have been written at least once since <code>&lt;date&gt;<\/code>) are left around.<\/p>\n<p>The idea is that objects which have been written recently are more likely to become reachable again in the future, and would thus be more likely to be susceptible to the kind of race we talked about above if they were to be pruned. Objects which haven&#8217;t been written recently, on the other hand, are proportionally less likely to become reachable again, and so they are safe (or, at least, safer) to remove.<\/p>\n<p>This idea isn&#8217;t foolproof, and it is certainly possible to run into the race we talked about earlier. We&#8217;ll discuss one such scenario towards the end of this post (along with the way we worked around it). But in practice, this strategy is simple and effective, preventing most instances of potential repository corruption.<\/p>\n<h2 id=\"storing-loose-unreachable-objects\"><a class=\"heading-link\" href=\"#storing-loose-unreachable-objects\">Storing loose unreachable objects<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>But one question remains: how does Git keep track of the age of unreachable objects which haven&#8217;t yet aged out of the repository?<\/p>\n<p>The answer, though simple, is at the heart of the problem we&#8217;re trying to solve here. Unreachable objects which have been written too recently to be removed from the repository are stored as loose objects, the individual object files stored in <code>.git\/objects<\/code>. Storing these unreachable objects individually means that we can rely on their <code><a href=\"https:\/\/man7.org\/linux\/man-pages\/man2\/lstat.2.html\">stat()<\/a><\/code> modification time (hereafter, <code>mtime<\/code>) to tell us how recently they were written.<\/p>\n<p>But this leads to an unfortunate problem: if a repository has many unreachable objects, and a large number of them were written recently, they must all be stored individually as loose objects. This is undesirable for a number of reasons:<\/p>\n<ul>\n<li>Pairs of unreachable objects that share a vast majority of their contents must be stored separately, and can&#8217;t benefit from the kind of deduplication offered by <a href=\"https:\/\/github.blog\/2022-08-29-gits-database-internals-i-packed-object-store\/\">packfiles<\/a>. This can cause your repository to take up much more space than it otherwise would.<\/li>\n<li>Having too many files (especially too many in a single directory) can lead to performance problems, including exhausting your system&#8217;s available <a href=\"https:\/\/en.wikipedia.org\/wiki\/Inode\">inodes<\/a> in the extreme case, leaving you unable to create new files, even if there may be space available for them.<\/li>\n<li>Any Git operation which has to scan through all loose objects (for example, <code>git repack -d<\/code>, which creates a new pack containing just your repository&#8217;s unpacked objects) will slow down as there are more files to process.<\/li>\n<\/ul>\n<p>It&#8217;s tempting to want to store all of a repository&#8217;s unreachable objects into a single pack. But there&#8217;s a problem there, too. Since all of the objects in a single pack share the same <code>mtime<\/code> (the <code>mtime<\/code> of the <code>*.pack<\/code> file itself), rewriting any single unreachable object has the effect of updating the <code>mtime<\/code>s of all of a repository&#8217;s unreachable objects. This is because Git optimizes out object writes for packed objects by simply updating the <code>mtime<\/code> of any pack(s) which contain that object. This makes it nearly impossible to expire any objects out of the repository permanently.<\/p>\n<h2 id=\"cruft-packs\"><a class=\"heading-link\" href=\"#cruft-packs\">Cruft packs<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>To solve this problem, we turned to a <a href=\"https:\/\/lore.kernel.org\/git\/20120611160824.GB12773@sigill.intra.peff.net\/\">long-discussed idea<\/a> on the <a href=\"https:\/\/lore.kernel.org\/git\">Git mailing list<\/a>: cruft packs. The idea is simple: store an auxiliary list of <code>mtime<\/code> data alongside a pack containing just unreachable objects. To garbage collect a repository, Git places the unreachable objects in a pack. That pack is designated as a &#8220;cruft pack&#8221; because Git also writes the <code>mtime<\/code> data corresponding to each object in a separate file alongside that pack. This makes it possible to update the <code>mtime<\/code> of a single unreachable object without changing the <code>mtime<\/code>s of any other unreachable object.<\/p>\n<p>To give you a sense of what this looks like in practice, here&#8217;s a small example:<\/p>\n<p><img data-recalc-dims=\"1\" decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=1024&#038;resize=1024%2C485\" alt=\"a pack of Git objects (represented by rectangles of different colors)\" width=\"1024\" height=\"485\" class=\"aligncenter size-large wp-image-67032 width-fit\" srcset=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=2468 2468w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=300 300w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=768 768w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=1024 1024w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=1536 1536w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?w=2048 2048w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>The above figure shows a pack of Git objects (represented by rectangles of different colors), its pack index, and the new <code>.mtimes<\/code> file. Together, these three files make up what Git calls a &#8220;<a href=\"https:\/\/git-scm.com\/docs\/cruft-packs\/2.37.0\">cruft pack<\/a>,\u201d and it&#8217;s what allows Git to store unreachable objects together, without needing a single file for each object.<\/p>\n<p>So, how do they work? Git uses the cruft pack to store a collection of object <code>mtime<\/code>s together in an array stored in the <code>*.mtimes<\/code> file. In order to discover the <code>mtime<\/code> for an individual object in a pack, Git first does a <a href=\"https:\/\/en.wikipedia.org\/wiki\/Binary_search_algorithm\">binary search<\/a> on the pack&#8217;s index to discover that object&#8217;s <a href=\"https:\/\/en.wikipedia.org\/wiki\/Lexicographic_order\">lexicographic index<\/a>. Git can then use that offset to read a 4-byte, unsigned integer in the <code>.mtimes<\/code> file. The <code>.mtimes<\/code> file contains a table of integers (one for each object in the associated <code>*.pack<\/code> file), each representing an <a href=\"https:\/\/en.wikipedia.org\/wiki\/Epoch_(computing)\">epoch timestamp<\/a> containing that object&#8217;s <code>mtime<\/code>. In other words, the <code>*.mtimes<\/code> file has a table of numbers, where each number represents an individual object&#8217;s <code>mtime<\/code>, encoded as a number of seconds since the <a href=\"https:\/\/en.wikipedia.org\/wiki\/Unix_epoch\">Unix epoch<\/a>.<\/p>\n<p>Crucially, this makes it possible to store all of a repository&#8217;s unreachable objects together in a single pack, without having to store them as individual loose objects, bypassing all of the drawbacks we discussed in the last section. Moreover, it allows Git to update the <code>mtime<\/code> of a single unreachable object, without inadvertently triggering the same update across all unreachable objects.<\/p>\n<p>Since Git doesn&#8217;t portably support updating a file in place, updating an object&#8217;s <code>mtime<\/code> (a process which Git calls &#8220;freshening&#8221;) takes place by writing a separate copy of that object out as a loose file. Of course, if we had to freshen all objects in a cruft pack, we would end up in a situation no better than before. But such updates tend to be unlikely in practice, and so writing individual copies of a small handful of unreachable objects ends up being a reasonable trade off most of the time.<\/p>\n<h2 id=\"generating-cruft-packs\"><a class=\"heading-link\" href=\"#generating-cruft-packs\">Generating cruft packs<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>Now that we have introduced the concept of cruft packs, the question remains: how does Git generate them?<\/p>\n<p>Despite being called <code>git gc<\/code> (short for &#8220;garbage collection&#8221;), running <code>git gc<\/code> does not always result in deleting unreachable objects. If you run <code>git gc --prune=never<\/code>, then Git will repack all reachable objects and move all unreachable objects to the cruft pack. If, however, you run <code>git gc --prune=1.day.ago<\/code>, then Git will repack all reachable objects, delete any unreachable objects that are older than one day, and repack the remaining unreachable objects into the cruft pack.<\/p>\n<p>This is because of Git&#8217;s treatment of unreachable parts of the repository. While Git only relies on having a reachability closure over reachable objects, Git&#8217;s garbage collection routine tries to leave unreachable parts of the repository intact to the extent possible. That means if Git encounters some unreachable cluster of objects in your repository, it will either expire all or none of those objects, but never some subset of them.<\/p>\n<p>We&#8217;ll discuss how cruft packs are generated with and without object expiration in the two sections below.<\/p>\n<h3 id=\"cruft-packs-without-object-expiration\"><a class=\"heading-link\" href=\"#cruft-packs-without-object-expiration\">Cruft packs without object expiration<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>When generating a cruft pack with an object expiration of <code>--date=never<\/code>, our only goal is to collect all unreachable objects together into a single cruft pack. Broadly speaking, this occurs in three steps:<\/p>\n<ol>\n<li>Starting at all of the branches and tags, generate a pack containing only reachable objects.<\/li>\n<li>Looking at all other existing packs, enumerate the list of objects which don&#8217;t appear in the new pack of reachable objects. Create a new pack containing just these objects, which are unreachable.<\/li>\n<li>Delete the existing packs.<\/li>\n<\/ol>\n<p>If any of that was confusing, don&#8217;t worry: we&#8217;ll break it down here step by step. The first step to collecting a repository&#8217;s unreachable objects is to figure out the parts of it that are reachable. If you&#8217;ve ever run <code>git repack -A<\/code>, this is exactly how that command works. Git starts a reachability traversal beginning at each of the branches and tags in your repository. Then it traverses back through history by walking from commits to their parents, trees to their sub-trees, and so on, marking every object that it sees along the way as reachable.<\/p>\n<p><img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection4.gif\" alt=\"Demonstration of how Git walks through a commit graph, from commit to parent\"><\/p>\n<p>Here, we&#8217;re showing the same commit graph from earlier in the post. Git&#8217;s goal at this point is simply to mark every reachable object that it sees, and it&#8217;s those objects that will become the contents of a new pack containing just reachable objects. Git starts by examining each reference, and walking from a commit to its parents until it either finds a commit with no parents (indicating the beginning of history), or a commit that it has already marked as reachable.<\/p>\n<p>In the above, the commit being walked is highlighted in dark blue, and any commits marked as reachable are marked in green. At each step, the commit currently being visited gets marked as reachable, and its parent(s) are visited in the next step. By repeating this process among all branches and tags, Git will mark all reachable objects in the repository.<\/p>\n<p>We can then use this set of objects to produce a new pack containing all reachable objects in a repository. Next, Git needs to discover the set of objects that it didn&#8217;t mark in the previous stage. A reasonable first approach might be to store the IDs of all of a repository&#8217;s objects in a set, and then remove them one by one as we mark objects reachable along our walk.<\/p>\n<p>But this approach tends to be impractical, since each object will require a minimum of 20 bytes of memory in order to insert into this set. At the time of writing, the linux.git repository contains nearly nine million objects, which would require nearly 180 MB of memory just to write out all of their object IDs.<\/p>\n<p>Instead, Git looks through all of the objects in all of the existing packs, checking whether or not each is contained in the new pack of reachable objects. Any object found in an existing pack which doesn\u2019t appear in the reachable pack is automatically included in the cruft pack.<\/p>\n<p><img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection5.gif\" alt=\"Animation demonstrating how  Git looks through all of the objects in all of the existing packs, checking whether or not each is contained in the new pack of reachable objects.\"><\/p>\n<p>Here, we&#8217;re going one by one among all of the pre-existing packs (here, labeled as <code>pack-abc.pack<\/code>, <code>pack-def.pack<\/code>, and <code>pack-123.pack<\/code>) and inspecting their objects one at a time. We first start with object <code>c8<\/code>, looking through the reachable pack (denoted as <code>pack-xyz.pack<\/code>) to see if any of its objects match <code>c8<\/code>. Since none do, <code>c8<\/code> is marked unreachable (which we represent by filling the object with a red background).<\/p>\n<p>This process is repeated for each object in each existing pack. Once this process is complete, all objects that existed in the repository before starting a garbage collection are marked either green, or red (indicating that they are either reachable, or unreachable, respectively).<\/p>\n<p>Git can then use the set of unreachable objects to generate a new pack, like below:<\/p>\n<p><img data-recalc-dims=\"1\" decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=1024&#038;resize=1024%2C380\" alt=\"A set of labeled Git packs\" width=\"1024\" height=\"380\" class=\"aligncenter size-large wp-image-67040 width-fit\" srcset=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=1600 1600w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=300 300w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=768 768w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=1024 1024w, https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection6-1.png?w=1536 1536w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>This pack (on the far right of the above image, denoted <code>pack-cruft.pack<\/code>) contains exactly the set of unreachable objects present in the repository at the beginning of garbage collection. By keeping track of each unreachable object&#8217;s <code>mtime<\/code> while marking existing objects, Git has enough data to write out a <code>*.mtimes<\/code> file in addition to the new pack, leaving us with a cruft pack containing just the repository&#8217;s unreachable objects.<\/p>\n<p>Here, we&#8217;re eliding some technical details about keeping track of each object&#8217;s <code>mtime<\/code> along the way, for brevity and simplicity. The routine is straightforward, though: each time we discover an object, we mark its <code>mtime<\/code> based on how we discovered the object.<\/p>\n<ul>\n<li>If an object is found in a packfile, it inherits its <code>mtime<\/code> from the packfile itself.<\/li>\n<li>If an object is found as a loose object, its <code>mtime<\/code> comes from the loose object file.<\/li>\n<li>And if an object is found in an existing cruft pack, its <code>mtime<\/code> comes from reading the cruft pack&#8217;s <code>*.mtimes<\/code> file at the appropriate index.<\/li>\n<\/ul>\n<p>If an object is seen more than once (e.g., an unreachable object stored in a cruft pack was freshened, resulting in another loose copy of the object), the <code>mtime<\/code> which is ultimately recorded in the new cruft pack is the most recent <code>mtime<\/code> of all of the above.<\/p>\n<h3 id=\"cruft-packs-with-object-expiration\"><a class=\"heading-link\" href=\"#cruft-packs-with-object-expiration\">Cruft packs with object expiration<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Generating cruft packs where some objects are going to expire out of the repository follows a similar, but slightly trickier approach than in the non-expiring case.<\/p>\n<p>Doing a garbage collection with a fixed expiration is known as &#8220;pruning.\u201d This essentially boils down to asking Git to pack the contents of a repository into two packfiles: one containing reachable objects, and another containing any unreachable objects. But, it also means that for some fixed expiration date, any unreachable objects which have an <code>mtime<\/code> older than the expiration date are removed from the repository entirely.<\/p>\n<p>The difficulty in this case stems from a fact briefly mentioned earlier in this post, which is that Git attempts to prevent connected clusters of unreachable objects from leaving the repository if some, but not all, of their objects have aged out.<\/p>\n<p>To make things clearer, here&#8217;s an example. Suppose that a repository has a handful of blob objects, all connected to some tree object, and all of these objects are unreachable. Assuming that they&#8217;re all old enough, then they will all expire together: no big deal. But what if the tree isn&#8217;t old enough to be expired? In this case, even though the blobs connected to it could be expired on their own, Git will keep them around since they&#8217;re connected to a tree with a sufficiently recent <code>mtime<\/code>. Git does this to preserve the repository&#8217;s reachability closure in case that tree were to become reachable again (in which case, having the tree and its blobs becomes important).<\/p>\n<p>To ensure that Git preserves any unreachable objects which are reachable from recent objects Git handles this case of cruft pack generation slightly differently. At a high level, it:<\/p>\n<ol>\n<li>Generates a candidate list of cruft objects, using the same process as outlined in the previous section.<\/li>\n<li>Then, to determine the actual list of cruft objects to keep around, it performs a reachability traversal using all of the candidate cruft objects, adding any object it sees along the way to the cruft pack.<\/li>\n<\/ol>\n<p>To make things a little clearer, here&#8217;s an example:<\/p>\n<p><img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection7.gif\" alt=\"Animation of Git performing  a reachability traversal\"><\/p>\n<p>After determining the set of unreachable objects (represented above as colored red) Git does a reachability traversal from each entry point into the graph of unreachable objects. Above, commits are represented by circles, trees by rectangles, and tree entries as rows within the larger rectangles. The <code>mtime<\/code>s are written below each commit.<\/p>\n<p>For now, let&#8217;s assume our expiration date is <code>d<\/code>, so any object whose <code>mtime<\/code> is greater than d must stay (despite being unreachable), and anything older than <code>d<\/code> can be pruned. Git traverses through each entry and asks, &#8220;Is this object old enough to be pruned?&#8221; When the answer is &#8220;yes\u201d Git leaves the object alone and moves on to the next entry point. When the answer is &#8220;no,\u201d however, (ie., Git is looking at an unreachable object whose <code>mtime<\/code> is too recent to prune), Git marks that object as &#8220;rescued&#8221; (indicated by turning it green) and then continues its traversal, marking any reachable objects as rescued.<\/p>\n<p>Objects that are rescued during this pass are written to the cruft pack, preserving their existence in the repository, leaving them to either continue to age, or have their <code>mtime<\/code>s updated before the next garbage collection.<\/p>\n<p>Let&#8217;s take a closer look at the example above. Git starts by looking at object <code>C<sub>(1,1)<\/sub><\/code>, and notice that its <code>mtime<\/code> is <code>d+5<\/code>, meaning that (since it happens after our expiration time, <code>d<\/code>) it is too new to expire. That causes Git to start a reachability traversal beginning at <code>C<sub>(1,1)<\/sub><\/code>, rescuing every object it encounters along the way. Since many objects are shared between multiple commits, rescuing an object from a more recent part of the graph often ends up marking older objects as rescued, too.<\/p>\n<p>After finishing the rescuing pass focused on <code>C<sub>(1,1)<\/sub><\/code>, Git moves on to look at <code>C<sub>(0,2)<\/sub><\/code>. But this commit&#8217;s <code>mtime<\/code> is <code>d-10<\/code>, which is before our expiration cutoff of <code>d<\/code>, meaning that it is safe to remove. Git can skip looking at any objects reachable from this commit, since none of them will be rescued.<\/p>\n<p>Finally, Git looks at another connected cluster of the unreachable object graph, beginning at <code>C<sub>(3,1)<\/sub><\/code>. Since this object has an <code>mtime<\/code> of <code>d+10<\/code>, it is too new to expire, so Git performs another reachability traversal, rescuing it and any objects reachable from it.<\/p>\n<p>Notice that in the final graph state that the main cluster of commits (the one beginning with <code>C<sub>(0,2)<\/sub><\/code>) is only partially rescued. In fact, only the objects necessary to retain a reachability closure over the rescued objects among that cluster are saved from being pruned. So even though, for example, commit <code>C<sub>(2,1)<\/sub><\/code> has only part of its tree entries rescued, that is OK since <code>C<sub>(2,1)<\/sub><\/code> itself will be pruned (hence any non-rescued tree entries connected to it are unimportant and will also be pruned).<\/p>\n<h2 id=\"putting-it-all-together\"><a class=\"heading-link\" href=\"#putting-it-all-together\">Putting it all together<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>Now that Git can generate a cruft pack and perform garbage collection on a repository with or without pruning objects, it was time to put all of the pieces together and submit the patches to the open-source Git project.<\/p>\n<p>Other Git sub-commands, like <code>repack<\/code>, and <code>gc<\/code> needed to learn about cruft packs, and gain command-line flags and configuration knobs in order to opt-in to the new behavior. With all of the pieces in place, you can now trigger a garbage collection by running either:<\/p>\n<pre><code>$ git gc --prune=1.day.ago --cruft\n<\/code><\/pre>\n<p>or<\/p>\n<pre><code>$ git repack -d --cruft --cruft-expiration=1.day.ago\n<\/code><\/pre>\n<p>to repack your repository into a reachable pack, and a cruft pack containing unreachable objects whose <code>mtime<\/code>s are within the past day. More details on the new command-line options and configuration can be found <a href=\"https:\/\/git-scm.com\/docs\/git-gc\/v2.37.0#Documentation\/git-gc.txt---cruft\">here<\/a>, <a href=\"https:\/\/git-scm.com\/docs\/git-repack\/v2.37.0#Documentation\/git-repack.txt---cruft\">here<\/a>, <a href=\"https:\/\/git-scm.com\/docs\/git-config\/v2.37.0#Documentation\/git-config.txt-gccruftPacks\">here<\/a>, and <a href=\"https:\/\/git-scm.com\/docs\/git-config\/v2.37.0#Documentation\/git-config.txt-repackcruftWindow\">here<\/a>.<\/p>\n<p>GitHub submitted the entirety of the patches that comprise cruft packs to the open-source Git project, and the results were <a href=\"https:\/\/github.blog\/2022-06-27-highlights-from-git-2-37\/#a-new-mechanism-for-pruning-unreachable-objects\">released in v2.37.0<\/a>. That means that you can use the same tools as what we run at GitHub on your own laptop, to run garbage collection on your own repositories.<\/p>\n<p>For those curious about the details, you can read the complete thread on the mailing list archive <a href=\"https:\/\/lore.kernel.org\/git\/cover.1638224692.git.me@ttaylorr.com\/\">here<\/a>.<\/p>\n<h2 id=\"cruft-packs-at-github\"><a class=\"heading-link\" href=\"#cruft-packs-at-github\">Cruft packs at GitHub<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>After a lengthy process of testing to ensure that using cruft packs was safe to carry out across all repositories on GitHub, we deployed and enabled the feature across all repositories. We kept a close eye on repositories with large numbers of unreachable objects, since the process of breaking any deltas between reachable and unreachable objects (since the two are now stored in separate packs, and object deltas cannot cross pack boundaries) can cause the initial cruft pack generation to take a long time. A small handful of repositories with many unreachable objects needed more time to generate their very first cruft pack. In those instances, we generated their cruft packs outside of our normal repository maintenance jobs to avoid triggering any timeouts.<\/p>\n<p>Now, every repository on GitHub and in GitHub Enterprise (in <a href=\"https:\/\/docs.github.com\/en\/enterprise-server@3.3\/admin\/release-notes\">version 3.3<\/a> and newer) uses cruft packs to store their unreachable objects. This has made garbage collecting repositories (especially busy ones with many unreachable objects) tractable where it often required significant human intervention before. Before cruft packs, many repositories which required clean up were simply out of our reach because of the possibility of creating an explosion of loose objects which could derail performance for all repositories stored on a fileserver. Now, garbage collecting a repository is a simple task, no matter its size or scale.<\/p>\n<p>During our testing, we ran garbage collection on a handful of repositories, and got some exciting results. For repositories that regularly force-push a single commit to their main branch (leaving a majority of their objects unreachable), their on-disk size dropped significantly. The most extreme example we found during testing caused a repository which used to take 186 gigabytes to store shrink to only take 2 gigabytes of space.<\/p>\n<p>On <code>github\/github<\/code>, GitHub&#8217;s main codebase, we were able to shrink the repository from around 57 gigabytes to 27 gigabytes. Even though these savings are more modest, the real payoff is in the objects we no longer have to store. Before garbage collecting, each replica of this repository had nearly 60 million objects, including years of test-merges, force-pushes, and all kinds of sources of unreachable objects. Each of these objects contributed to the I\/O cost of repacking this repository. After garbage collecting, only 11.8 million objects remained. Since each object in a repository requires around 150 bytes of memory during repacking, we save around 7 gigabytes of RAM during each maintenance routine.<\/p>\n<h2 id=\"limbo-repositories\"><a class=\"heading-link\" href=\"#limbo-repositories\">Limbo repositories<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>Even though we can easily garbage collect a repository of any size, we still have to navigate the inherent raciness that we described at the beginning of this post.<\/p>\n<p>At GitHub, our approach has been to make this situation easy to recover from automatically instead of preventing it entirely (which would require significant surgery to much of Git&#8217;s code). To do this, our approach is to create a &#8220;limbo&#8221; repository whenever a pruning garbage collection is done. Any objects which get expired from the main repository are stored in a separate pack in the limbo repository. Then, the process to garbage collect a repository looks something like:<\/p>\n<ol>\n<li>Generate a cruft pack of recent unreachable objects in the main repository.<\/li>\n<li>Generate a second cruft pack of expired unreachable objects, stored outside of the main repository, in the &#8220;limbo&#8221; repository.<\/li>\n<li>After garbage collection has completed, run a <code>git fsck<\/code> in the main repository to detect any object corruption.<\/li>\n<li>If any objects are missing, recover them by copying them over from the limbo repository.<\/li>\n<\/ol>\n<p>The process for generating a cruft pack of expired unreachable objects boils down to creating another cruft pack (using exactly the same process we described earlier in this post), with two caveats:<\/p>\n<ul>\n<li>The expiration cutoff is set to &#8220;never&#8221; since we want to keep around any objects which we did expire in the previous step.<\/li>\n<li>The original cruft pack is treated as a pack containing reachable objects since we want to ignore any unreachable objects which were too recent to expire (and, thus, are stored in the cruft pack in the main repository).<\/li>\n<\/ul>\n<p>We have used this idea at GitHub with great success, and now treat garbage collection as a hands-off process from start to finish. The patches to implement this approach are available as a preliminary RFC on the Git mailing list <a href=\"https:\/\/lore.kernel.org\/git\/cover.1656528343.git.me@ttaylorr.com\/\">here<\/a>.<\/p>\n<h2 id=\"thank-you\"><a class=\"heading-link\" href=\"#thank-you\">Thank you<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>This work would not have been possible without generous review and collaboration from engineers from within and outside of GitHub. The Git Systems team at GitHub were great to work with while we developed and deployed cruft packs. Special thanks to <a href=\"https:\/\/github.com\/torstenwalter\">Torsten Walter<\/a>, and <a href=\"https:\/\/github.com\/mhagger\">Michael Haggerty<\/a>, who played substantial roles in developing limbo repositories.<\/p>\n<p>Outside of GitHub, this work would not have been possible without careful review from the open-source Git community, especially <a href=\"https:\/\/github.com\/derrickstolee\">Derrick Stolee<\/a>, <a href=\"https:\/\/github.com\/peff\">Jeff King<\/a>, <a href=\"https:\/\/github.com\/jonathantanmy\">Jonathan Tan<\/a>, <a href=\"https:\/\/github.com\/jrn\">Jonathan Nieder<\/a>, and <a href=\"https:\/\/github.com\/gitster\">Junio C Hamano<\/a>. In particular, Jeff King contributed significantly to the original development of many of the ideas discussed above.<\/p>\n<p><!-- Footnotes themselves at the bottom. --><\/p>\n<h2 id=\"notes\"><a class=\"heading-link\" href=\"#notes\">Notes<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<div class=\"footnotes\">\n<hr \/>\n<ol>\n<li id=\"fn-67031-1\">\nIt&#8217;s true. According to the Library of Congress themselves, their digital collection amounts to more than 3 petabytes in size [<a href=\"https:\/\/blogs.loc.gov\/thesignal\/2012\/04\/a-library-of-congress-worth-of-data-its-all-in-how-you-define-it\/\">source<\/a>]. The 18.6 petabytes we store at GitHub actually overcounts by a factor of five, since we store a handful of copies of each repository. In reality, it&#8217;s hard to provide an exact number, since data is de-duplicated within a fork network, and is stored compressed on disk. Either way you slice it, it&#8217;s a lot of data: you get the point.&#160;<a href=\"#fnref-67031-1\" title=\"Return to main content.\">&#8617;<\/a>\n<\/li>\n<li id=\"fn-67031-2\">\nMeaning that for any reachable object part of some repository, any objects reachable from it are also contained in that repository.&#160;<a href=\"#fnref-67031-2\" title=\"Return to main content.\">&#8617;<\/a>\n<\/li>\n<\/ol>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.<\/p>\n","protected":false},"author":1282,"featured_media":67056,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_gh_post_show_toc":"no","_gh_post_is_no_robots":"","_gh_post_is_featured":"no","_gh_post_is_excluded":"no","_gh_post_is_unlisted":"","_gh_post_related_link_1":"","_gh_post_related_link_2":"","_gh_post_related_link_3":"","_gh_post_sq_img":"https:\/\/github.blog\/wp-content\/uploads\/2022\/01\/git-thumbnail.png","_gh_post_sq_img_id":"62621","_gh_post_cta_title":"","_gh_post_cta_text":"","_gh_post_cta_link":"","_gh_post_cta_button":"Click Here to Learn More","_gh_post_recirc_hide":"no","_gh_post_recirc_col_1":"gh-auto-select","_gh_post_recirc_col_2":"66717","_gh_post_recirc_col_3":"65308","_gh_post_recirc_col_4":"65316","_featured_video":"","_gh_post_additional_query_params":"","_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"{title}\n\n{excerpt}\n\n{url}","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"_wpas_customize_per_network":false,"jetpack_post_was_ever_published":false,"_links_to":"","_links_to_target":""},"categories":[3307,72],"tags":[132],"coauthors":[2189],"class_list":["post-67031","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-architecture-optimization","category-engineering","tag-git"],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.4 (Yoast SEO v28.4) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Scaling Git\u2019s garbage collection - The GitHub Blog<\/title>\n<meta name=\"description\" content=\"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Scaling Git\u2019s garbage collection\" \/>\n<meta property=\"og:description\" content=\"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/\" \/>\n<meta property=\"og:site_name\" content=\"The GitHub Blog\" \/>\n<meta property=\"article:published_time\" content=\"2022-09-13T16:00:49+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2022-09-13T16:54:52+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?fit=2468%2C1170\" \/>\n\t<meta property=\"og:image:width\" content=\"2468\" \/>\n\t<meta property=\"og:image:height\" content=\"1170\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Taylor Blau\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?fit=2468%2C1170\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Taylor Blau\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"24 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/\"},\"author\":{\"name\":\"Taylor Blau\",\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/f2a5dc09d09f41c8c731679cc07da524\"},\"headline\":\"Scaling Git\u2019s garbage collection\",\"datePublished\":\"2022-09-13T16:00:49+00:00\",\"dateModified\":\"2022-09-13T16:54:52+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/\"},\"wordCount\":4945,\"image\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2022\\\/09\\\/Untitled.png?fit=2400%2C1260\",\"keywords\":[\"Git\"],\"articleSection\":[\"Architecture &amp; optimization\",\"Engineering\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/\",\"url\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/\",\"name\":\"Scaling Git\u2019s garbage collection - The GitHub Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/github.blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2022\\\/09\\\/Untitled.png?fit=2400%2C1260\",\"datePublished\":\"2022-09-13T16:00:49+00:00\",\"dateModified\":\"2022-09-13T16:54:52+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/f2a5dc09d09f41c8c731679cc07da524\"},\"description\":\"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#primaryimage\",\"url\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2022\\\/09\\\/Untitled.png?fit=2400%2C1260\",\"contentUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2022\\\/09\\\/Untitled.png?fit=2400%2C1260\",\"width\":2400,\"height\":1260},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/scaling-gits-garbage-collection\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/github.blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Engineering\",\"item\":\"https:\\\/\\\/github.blog\\\/engineering\\\/\"},{\"@type\":\"ListItem\",\"position\":3,\"name\":\"Architecture &amp; optimization\",\"item\":\"https:\\\/\\\/github.blog\\\/engineering\\\/architecture-optimization\\\/\"},{\"@type\":\"ListItem\",\"position\":4,\"name\":\"Scaling Git\u2019s garbage collection\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/github.blog\\\/#website\",\"url\":\"https:\\\/\\\/github.blog\\\/\",\"name\":\"The GitHub Blog\",\"description\":\"Updates, ideas, and inspiration from GitHub to help developers build and design software.\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/github.blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/f2a5dc09d09f41c8c731679cc07da524\",\"name\":\"Taylor Blau\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g2ce44289191883c54a58a554d8fc874a\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g\",\"caption\":\"Taylor Blau\"},\"description\":\"Taylor Blau is a Principal Software Engineer at GitHub where he works on Git.\",\"sameAs\":[\"https:\\\/\\\/ttaylorr.com\"],\"url\":\"https:\\\/\\\/github.blog\\\/author\\\/ttaylorr\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Scaling Git\u2019s garbage collection - The GitHub Blog","description":"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/","og_locale":"en_US","og_type":"article","og_title":"Scaling Git\u2019s garbage collection","og_description":"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.","og_url":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/","og_site_name":"The GitHub Blog","article_published_time":"2022-09-13T16:00:49+00:00","article_modified_time":"2022-09-13T16:54:52+00:00","og_image":[{"width":2468,"height":1170,"url":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?fit=2468%2C1170","type":"image\/png"}],"author":"Taylor Blau","twitter_card":"summary_large_image","twitter_image":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/garbagecollection1.png?fit=2468%2C1170","twitter_misc":{"Written by":"Taylor Blau","Est. reading time":"24 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#article","isPartOf":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/"},"author":{"name":"Taylor Blau","@id":"https:\/\/github.blog\/#\/schema\/person\/f2a5dc09d09f41c8c731679cc07da524"},"headline":"Scaling Git\u2019s garbage collection","datePublished":"2022-09-13T16:00:49+00:00","dateModified":"2022-09-13T16:54:52+00:00","mainEntityOfPage":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/"},"wordCount":4945,"image":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#primaryimage"},"thumbnailUrl":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/Untitled.png?fit=2400%2C1260","keywords":["Git"],"articleSection":["Architecture &amp; optimization","Engineering"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/","url":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/","name":"Scaling Git\u2019s garbage collection - The GitHub Blog","isPartOf":{"@id":"https:\/\/github.blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#primaryimage"},"image":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#primaryimage"},"thumbnailUrl":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/Untitled.png?fit=2400%2C1260","datePublished":"2022-09-13T16:00:49+00:00","dateModified":"2022-09-13T16:54:52+00:00","author":{"@id":"https:\/\/github.blog\/#\/schema\/person\/f2a5dc09d09f41c8c731679cc07da524"},"description":"A tour of recent work to re-engineer Git\u2019s garbage collection process to scale to our largest and most active repositories.","breadcrumb":{"@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#primaryimage","url":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/Untitled.png?fit=2400%2C1260","contentUrl":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/Untitled.png?fit=2400%2C1260","width":2400,"height":1260},{"@type":"BreadcrumbList","@id":"https:\/\/github.blog\/engineering\/architecture-optimization\/scaling-gits-garbage-collection\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/github.blog\/"},{"@type":"ListItem","position":2,"name":"Engineering","item":"https:\/\/github.blog\/engineering\/"},{"@type":"ListItem","position":3,"name":"Architecture &amp; optimization","item":"https:\/\/github.blog\/engineering\/architecture-optimization\/"},{"@type":"ListItem","position":4,"name":"Scaling Git\u2019s garbage collection"}]},{"@type":"WebSite","@id":"https:\/\/github.blog\/#website","url":"https:\/\/github.blog\/","name":"The GitHub Blog","description":"Updates, ideas, and inspiration from GitHub to help developers build and design software.","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/github.blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/github.blog\/#\/schema\/person\/f2a5dc09d09f41c8c731679cc07da524","name":"Taylor Blau","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g2ce44289191883c54a58a554d8fc874a","url":"https:\/\/secure.gravatar.com\/avatar\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/d5f3476f26b6f99cbb6b467e7ed7482f5762c8157bc73f569196e428bdcbea25?s=96&d=mm&r=g","caption":"Taylor Blau"},"description":"Taylor Blau is a Principal Software Engineer at GitHub where he works on Git.","sameAs":["https:\/\/ttaylorr.com"],"url":"https:\/\/github.blog\/author\/ttaylorr\/"}]}},"jetpack_publicize_connections":[],"jetpack_shortlink":"https:\/\/wp.me\/pamS32-hr9","jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/github.blog\/wp-content\/uploads\/2022\/09\/Untitled.png?fit=2400%2C1260","_links":{"self":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/67031","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/users\/1282"}],"replies":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/comments?post=67031"}],"version-history":[{"count":9,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/67031\/revisions"}],"predecessor-version":[{"id":67034,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/67031\/revisions\/67034"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/media\/67056"}],"wp:attachment":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/media?parent=67031"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/categories?post=67031"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/tags?post=67031"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/coauthors?post=67031"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}