{"id":60371,"date":"2021-09-27T09:02:13","date_gmt":"2021-09-27T16:02:13","guid":{"rendered":"https:\/\/github.blog\/?p=60371"},"modified":"2022-03-02T09:44:58","modified_gmt":"2022-03-02T17:44:58","slug":"partitioning-githubs-relational-databases-scale","status":"publish","type":"post","link":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/","title":{"rendered":"Partitioning GitHub\u2019s relational databases to handle scale"},"content":{"rendered":"<p>More than 10 years ago, GitHub.com started out like many other web applications of that time\u2014built on Ruby on Rails, with a single MySQL database to store most of its data.<\/p>\n<p>Over the years, this architecture went through many iterations to support GitHub&#8217;s growth and ever-evolving resiliency requirements. For example, we started storing data for some features (like <a href=\"https:\/\/docs.github.com\/en\/rest\/reference\/repos#statuses\" rel=\"noopener\" target=\"_blank\">statuses<\/a>) in separate MySQL databases, we added read replicas to spread the load across multiple machines, and we started using ProxySQL to reduce the number of connections opened against our primary MySQL instances.<\/p>\n<p>Yet at its core, GitHub.com remained built around one main database cluster (called <code>mysql1<\/code>) that housed a large portion of the data used by core GitHub features, like user profiles, repositories, issues, and pull requests.<\/p>\n<p>With GitHub\u2019s growth, this inevitably led to challenges. We struggled to keep our database systems adequately sized, always moving to newer and bigger machines to scale up. Any sort of incident negatively affecting <code>mysql1<\/code> would affect all features that stored their data on this cluster.<\/p>\n<p>In 2019, in order to meet the growth and availability challenges we faced, we set a plan in motion to improve our tooling and our ability to partition relational databases. As you can imagine, this was a complex challenge necessitating the introduction and creation of various tools as outlined below.<\/p>\n<p>The result, we see in 2021, is a 50% load reduction on database hosts housing the data that once was on <code>mysql1<\/code>. This contributed significantly to reducing the number of database-related incidents and improved GitHub.com\u2019s reliability for all our users.<\/p>\n<h2 id=\"virtual-partitions\"><a class=\"heading-link\" href=\"#virtual-partitions\">Virtual partitions<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>The first concept we introduced was virtual partitions of database schemas. Before database tables can be moved physically, we have to make sure they are separated <em>virtually<\/em> in the application layer, and this has to happen without impacting teams working on new or existing features.<\/p>\n<p>To do that, we group database tables that belong together into schema domains and enforce boundaries between the domains with SQL linters. This allows us to safely partition data later without ending up with queries and transactions that span partitions.<\/p>\n<h3 id=\"schema-domains\"><a class=\"heading-link\" href=\"#schema-domains\">Schema domains<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Schema domains are a tool we came up with to implement virtual partitions. A schema domain describes a tightly coupled set of database tables that are frequently used together in queries (such as, when using table joins or subqueries) and transactions. For example, the <code>gists<\/code> schema domain contains all of the tables supporting the GitHub Gist feature&#8211;lik\u200be the <code>gists<\/code>, <code>gist_comments<\/code> and <code>starred_gists<\/code> tables. Since they belong together, they should stay together. A schema domain is the first step to codify that.<\/p>\n<p>Schema domains put clear boundaries in place and expose sometimes-hidden dependencies between features. In the Rails application, the information is stored in a simple YAML configuration located at <code>db\/schema-domains.yml<\/code>. Here\u2019s an example illustrating the contents of that file:<\/p>\n<pre><code class=\"language-yaml\">gists:\n  - gist_comments\n  - gists\n  - starred_gists\nrepositories:\n  - issues\n  - pull_requests\n  - repositories\nusers:\n  - avatars\n  - gpg_keys\n  - public_keys\n  - users\n<\/code><\/pre>\n<p>A linter makes sure that the list of tables in this file matches our database schema. In turn, the same linter enforces the assignment of a schema domain to every table.<\/p>\n<h3 id=\"sql-linters\"><a class=\"heading-link\" href=\"#sql-linters\">SQL linters<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Building on top of schema domains, two new SQL linters enforce virtual boundaries between domains. They identify any violating queries and transactions that span schema domains by adding a query annotation and treating them as exemptions. If a domain has no violations, it is virtually partitioned and ready to be physically moved to another database cluster.<\/p>\n<h4 id=\"query-linter\"><a class=\"heading-link\" href=\"#query-linter\">Query linter<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h4>\n<p>The query linter verifies that only tables belonging to the same schema domain can be referenced in the same database query. If it detects tables from different domains, it throws an exception with a helpful message for the developer to avoid the issue.<\/p>\n<p>Since the linter is only enabled in development and test environments, developers encounter violation errors early in the development process. In addition, during CI runs, the linter ensures that no new violations are introduced by accident.<\/p>\n<p>The linter has a way to suppress the exception by annotating the SQL query with a special comment: <code>\/* cross-schema-domain-query-exempted *\/<\/code><\/p>\n<p>We even <a href=\"https:\/\/github.com\/rails\/rails\/pull\/35617\" rel=\"noopener\" target=\"_blank\">built and upstreamed a new method to ActiveRecord<\/a> to make adding such a comment easier:<\/p>\n<pre><code class=\"language-ruby\">Repository.joins(:owner).annotate(\"cross-schema-domain-query-exempted\")\n# =&gt; SELECT * FROM `repositories` INNER JOIN `users` ON `users`.`id` = `repositories.owner_id` \/* cross-schema-domain-query-exempted *\/\n<\/code><\/pre>\n<p>By annotating all queries that cause failures, a backlog of queries needing modification can be built. Here&#8217;s a couple approaches we often use to eliminate exemptions:<\/p>\n<ol>\n<li>Sometimes, an exemption can easily be addressed by triggering separate queries instead of joining tables. One example is using <code>ActiveRecord<\/code>&#8216;s <code><a href=\"https:\/\/api.rubyonrails.org\/classes\/ActiveRecord\/QueryMethods.html#method-i-preload\" rel=\"noopener\" target=\"_blank\">preload<\/a><\/code> method instead of <code><a href=\"https:\/\/api.rubyonrails.org\/classes\/ActiveRecord\/QueryMethods.html#method-i-includes\" rel=\"noopener\" target=\"_blank\">includes<\/a><\/code>.\n<p>Another challenge is <code>has_many :through<\/code> relations that lead to <code>JOIN<\/code>s across tables from different schema domains. For that, we worked on a <a href=\"https:\/\/github.blog\/2021-07-12-adding-support-cross-cluster-associations-rails-7\/\" rel=\"noopener\" target=\"_blank\">generic solution that got upstreamed to Rails as well<\/a>: <code>has_many<\/code> now has a <code>\u200bdisable_joins<\/code>\u200b option that tells Active Record not to do any <code>JOIN<\/code> queries across the underlying tables. Instead, it runs several queries passing primary key values.<\/p>\n<\/li>\n<li>\n<p>Joining data in the application instead of in the database is another common solution. For example, replacing <code>INNER JOIN<\/code> statements with two separate queries and instead performing the &#8220;union&#8221; operation in Ruby (for example, <code>A.pluck(:b_id) &amp; B.where(id: ...)<\/code>).<\/p>\n<p>In some cases, this leads to surprising performance <em>improvements<\/em>. Depending on the data structure and cardinality, MySQL&#8217;s query planner can sometimes create suboptimal query execution plans, whereas an application-side join has a more stable performance cost.<\/p>\n<\/li>\n<\/ol>\n<p>As with almost all reliability and performance-related changes, we ship them behind <a href=\"https:\/\/github.com\/github\/scientist\" rel=\"noopener\" target=\"_blank\">Scientist experiments<\/a> that execute both the old and new implementations for a subset of requests, allowing us to assess the performance impact of each change.<\/p>\n<h4 id=\"transaction-linter\"><a class=\"heading-link\" href=\"#transaction-linter\">Transaction linter<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h4>\n<p>In addition to queries, transactions are a concern as well. Existing application code was written with a certain database schema in mind. MySQL transactions guarantee consistency across tables within a database. If a transaction includes queries to tables that will move to separate databases, it will no longer be able to guarantee consistency.<\/p>\n<p>To understand which transactions need to be reviewed, we introduced a transaction linter. Similar to the query linter, it verifies that all tables which are used together in a given transaction belong to the same schema domain.<\/p>\n<p>This linter runs in production with heavy sampling to keep the performance impact at a minimum. The linting results are collected and analyzed to understand where most cross-domain transactions happen, allowing us to decide to either update certain code paths or adapt our data model.<\/p>\n<p>In cases where transactional consistency guarantees are crucial, we extract data into new tables that belong to the same schema domain. This ensures they stay on the same database cluster and therefore continue to have transactional consistency. This often happens with <em>polymorphic tables<\/em> that house data from different schema domains (for example, a <code>reactions<\/code> table storing records for different features like issues, pull requests, discussions, etc.)<\/p>\n<h2 id=\"moving-data-without-downtime\"><a class=\"heading-link\" href=\"#moving-data-without-downtime\">Moving data without downtime<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>A schema domain that is virtually isolated is ready to be physically moved to another database cluster. To move tables on the fly, we use two different approaches: Vitess, and a custom write-cutover process.<\/p>\n<h3 id=\"vitess\"><a class=\"heading-link\" href=\"#vitess\">Vitess<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p><a href=\"https:\/\/vitess.io\" rel=\"noopener\" target=\"_blank\">Vitess<\/a> is a scaling layer on top of MySQL that helps with sharding needs. We use its <a href=\"https:\/\/vitess.io\/docs\/reference\/features\/sharding\/#supported-operations\" rel=\"noopener\" target=\"_blank\">vertical sharding feature<\/a> to move sets of tables together in production without downtime.<\/p>\n<p>To do that, we deploy Vitess&#8217; <a href=\"https:\/\/vitess.io\/docs\/reference\/programs\/vtgate\/\" rel=\"noopener\" target=\"_blank\">VTGate<\/a> in our Kubernetes clusters. These VTGate processes become the endpoint for the application to connect to, instead of direct connections to MySQL. They implement the same MySQL protocol and are indistinguishable from the application side.<\/p>\n<p>The VTGate processes know the current state of the Vitess setup and talk to the MySQL instances via another Vitess component: <a href=\"https:\/\/vitess.io\/docs\/reference\/programs\/vttablet\/\" rel=\"noopener\" target=\"_blank\">VTTablet<\/a>. Behind the scenes, Vitess&#8217; table moving feature is powered by <a href=\"https:\/\/vitess.io\/docs\/reference\/vreplication\/\" rel=\"noopener\" target=\"_blank\">VReplication<\/a>, which replicates data between database clusters.<\/p>\n<h3 id=\"write-cutover-process\"><a class=\"heading-link\" href=\"#write-cutover-process\">Write-cutover process<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>Because Vitess\u2019 adoption was still in its early stages at the beginning of 2020, we developed an alternative approach to move large sets of tables at once. This mitigated the risk of relying on a single solution to ensure the continued availability of GitHub.com.<\/p>\n<p>We use MySQL&#8217;s regular replication feature to feed data to another cluster. Initially, the new cluster is added to the replication tree of the old cluster. Then a script quickly executes a series of changes to effect the cutover.<\/p>\n<figure id=\"attachment_60373\"  class=\"wp-caption aligncenter mx-0\"><img data-recalc-dims=\"1\" decoding=\"async\" width=\"607\" height=\"412\" loading=\"lazy\" src=\"https:\/\/github.blog\/wp-content\/uploads\/2021\/09\/GitHub-MySQL-database-cluster-setup.png?resize=607%2C412\" alt=\"\" class=\"width-fit size-full wp-image-60373\" srcset=\"https:\/\/github.blog\/wp-content\/uploads\/2021\/09\/GitHub-MySQL-database-cluster-setup.png?w=607 607w, https:\/\/github.blog\/wp-content\/uploads\/2021\/09\/GitHub-MySQL-database-cluster-setup.png?w=300 300w\" sizes=\"auto, (max-width: 607px) 100vw, 607px\" \/><figcaption class=\"text-mono color-fg-muted mt-14px f5-mktg\">The MySQL database cluster setup before executing the write-cutover process<\/figcaption><\/figure>\n<p>Before running the script, we prepare the application and database replication so that a destination cluster called <code>cluster_b<\/code> is a sub-cluster of the existing <code>cluster_a<\/code>. <a href=\"https:\/\/github.com\/sysown\/proxysql\" rel=\"noopener\" target=\"_blank\">ProxySQL<\/a> is used for <a href=\"https:\/\/proxysql.com\/documentation\/multiplexing\/\" rel=\"noopener\" target=\"_blank\">multiplexing client connections<\/a> to MySQL primaries. The ProxySQL instance on <code>cluster_b<\/code> is configured to route all traffic to the <code>cluster_a<\/code> primary. The use of ProxySQL allows us to change database traffic routing quickly and with minimal impact on database clients\u2014in our case, the Rails application.<\/p>\n<p>With this setup, we can move database connections to <code>cluster_b<\/code>  without splitting anything. All read traffic still goes to hosts replicating from the <code>cluster_a<\/code> primary. All write traffic remains with the <code>cluster_a<\/code> primary too.<\/p>\n<p>In this situation, we run a cutover script executing the following:<\/p>\n<ol>\n<li>Enable read-only mode for the <code>cluster_a<\/code> primary. At this point, all writes to <code>cluster_a<\/code> and <code>cluster_b<\/code> are prevented. All web requests that try to write to these database primaries fail and result in 500s.<\/li>\n<li>Read the last executed <a href=\"https:\/\/dev.mysql.com\/doc\/refman\/5.7\/en\/replication-gtids-concepts.html\" rel=\"noopener\" target=\"_blank\">MySQL GTID<\/a> from the <code>cluster_a<\/code> primary.<\/li>\n<li>Poll the <code>cluster_b<\/code> primary to verify the last executed GTID has arrived.<\/li>\n<li>Stop replication on the <code>cluster_b<\/code> primary from  <code>cluster_a<\/code>.<\/li>\n<li>Update ProxySQL routing configuration on <code>cluster_b<\/code> to direct traffic to the <code>cluster_b<\/code> primary.<\/li>\n<li>Disable read-only mode for the <code>cluster_a<\/code> and <code>cluster_b<\/code> primaries.<\/li>\n<li>Celebrate!<\/li>\n<\/ol>\n<p>After thorough preparation and exercising, we learned that these six steps execute in only a few tens of milliseconds for our busiest database tables. Since we execute such cutovers during our lowest traffic time of day, we only cause a handful of user-facing errors because of failed writes. The results of this approach were better than we expected.<\/p>\n<h3 id=\"learnings\"><a class=\"heading-link\" href=\"#learnings\">Learnings<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h3>\n<p>The write-cutover process was used to split up <code>mysql1<\/code>, our original main database cluster. We moved 130 of our busiest tables at once\u2013\u2013those that power GitHub&#8217;s core features: repositories, issues, and pull requests. This process was created as a risk-mitigation strategy to have multiple, independent tools at our disposal. In addition, because of factors like deployment topology and read-your-writes support, we didn\u2019t choose Vitess as the tool to move database tables in every case. We anticipate the opportunity to use it for the majority of data migrations in the future though.<\/p>\n<h2 id=\"results\"><a class=\"heading-link\" href=\"#results\">Results<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>The main database cluster <code>mysql1<\/code>, mentioned in the introduction, housed a large portion of the data used by many of GitHub&#8217;s most important features, like users, repositories, issues, and pull requests. Since 2019 we achieved the ability to scale this relational database with the following results:<\/p>\n<ul>\n<li>In 2019, <code>mysql1<\/code> answered 950,000 queries\/s on average, 900,000 queries\/s on replicas, and 50,000 queries\/s on the primary.<\/li>\n<li>Today, in 2021, the same database tables are spread across several clusters. In two years, they saw continued growth, accelerating year-over-year. All hosts of these clusters combined answer 1,200,000 queries\/s on average (1,125,000 queries\/s on replicas, 75,000 queries\/s on the primaries). At the same time, the average load on each host halved.<\/li>\n<\/ul>\n<p>The load reduction contributed significantly to reducing the number of database-related incidents and improved GitHub.com\u2019s reliability for all our users.<\/p>\n<h2 id=\"more-partitioning\"><a class=\"heading-link\" href=\"#more-partitioning\">More partitioning<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>In addition to vertical partitioning to move database tables, we also use horizontal partitioning (aka sharding). This allows us to split database tables across multiple clusters, enabling more sustainable growth. We\u2019ll detail the tooling, linters, and Rails improvements related to this in a future blog post.<\/p>\n<h2 id=\"conclusion\"><a class=\"heading-link\" href=\"#conclusion\">Conclusion<span class=\"heading-hash pl-2 text-italic text-bold\" aria-hidden=\"true\"><\/span><\/a><\/h2>\n<p>Over the last 10 years, GitHub has been learning to scale according to its needs. We often choose to leverage \u201cboring\u201d technology that has been proven to work at our scale, as reliability remains the primary concern. But the combination of industry-proven tools with simple changes to our production code and its dependencies has provided us with a path for the continued growth of our databases into the future.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In 2019, to meet GitHub&#8217;s growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases. <\/p>\n","protected":false},"author":1491,"featured_media":57452,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_gh_post_show_toc":"no","_gh_post_is_no_robots":"","_gh_post_is_featured":"no","_gh_post_is_excluded":"no","_gh_post_is_unlisted":"","_gh_post_related_link_1":"","_gh_post_related_link_2":"","_gh_post_related_link_3":"","_gh_post_sq_img":"","_gh_post_sq_img_id":"","_gh_post_cta_title":"","_gh_post_cta_text":"","_gh_post_cta_link":"","_gh_post_cta_button":"Click Here to Learn More","_gh_post_recirc_hide":"","_gh_post_recirc_col_1":"","_gh_post_recirc_col_2":"","_gh_post_recirc_col_3":"","_gh_post_recirc_col_4":"","_featured_video":"","_gh_post_additional_query_params":"","_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"{title}\n\n{excerpt}\n\n{url}","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"_wpas_customize_per_network":false,"jetpack_post_was_ever_published":false,"_links_to":"","_links_to_target":""},"categories":[72,3309],"tags":[],"coauthors":[2001],"class_list":["post-60371","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engineering","category-infrastructure"],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.4 (Yoast SEO v28.4) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Partitioning GitHub\u2019s relational databases to handle scale - The GitHub Blog<\/title>\n<meta name=\"description\" content=\"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Partitioning GitHub\u2019s relational databases to handle scale\" \/>\n<meta property=\"og:description\" content=\"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/\" \/>\n<meta property=\"og:site_name\" content=\"The GitHub Blog\" \/>\n<meta property=\"article:published_time\" content=\"2021-09-27T16:02:13+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2022-03-02T17:44:58+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"630\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Thomas Maurer\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Thomas Maurer\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"10 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/\"},\"author\":{\"name\":\"Thomas Maurer\",\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/7abf5c5c1cb04b26e5a04536b62c4501\"},\"headline\":\"Partitioning GitHub\u2019s relational databases to handle scale\",\"datePublished\":\"2021-09-27T16:02:13+00:00\",\"dateModified\":\"2022-03-02T17:44:58+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/\"},\"wordCount\":1938,\"image\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2021\\\/04\\\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\",\"articleSection\":[\"Engineering\",\"Infrastructure\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/\",\"url\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/\",\"name\":\"Partitioning GitHub\u2019s relational databases to handle scale - The GitHub Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/github.blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2021\\\/04\\\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\",\"datePublished\":\"2021-09-27T16:02:13+00:00\",\"dateModified\":\"2022-03-02T17:44:58+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/7abf5c5c1cb04b26e5a04536b62c4501\"},\"description\":\"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#primaryimage\",\"url\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2021\\\/04\\\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\",\"contentUrl\":\"https:\\\/\\\/github.blog\\\/wp-content\\\/uploads\\\/2021\\\/04\\\/Blog_ENGINEERING_for-social.png?fit=1200%2C630\",\"width\":1200,\"height\":630},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/partitioning-githubs-relational-databases-scale\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/github.blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Engineering\",\"item\":\"https:\\\/\\\/github.blog\\\/engineering\\\/\"},{\"@type\":\"ListItem\",\"position\":3,\"name\":\"Infrastructure\",\"item\":\"https:\\\/\\\/github.blog\\\/engineering\\\/infrastructure\\\/\"},{\"@type\":\"ListItem\",\"position\":4,\"name\":\"Partitioning GitHub\u2019s relational databases to handle scale\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/github.blog\\\/#website\",\"url\":\"https:\\\/\\\/github.blog\\\/\",\"name\":\"The GitHub Blog\",\"description\":\"Updates, ideas, and inspiration from GitHub to help developers build and design software.\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/github.blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/github.blog\\\/#\\\/schema\\\/person\\\/7abf5c5c1cb04b26e5a04536b62c4501\",\"name\":\"Thomas Maurer\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g438a2a1efba5fc2060e7180d2732a58e\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g\",\"caption\":\"Thomas Maurer\"},\"url\":\"https:\\\/\\\/github.blog\\\/author\\\/tma\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Partitioning GitHub\u2019s relational databases to handle scale - The GitHub Blog","description":"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/","og_locale":"en_US","og_type":"article","og_title":"Partitioning GitHub\u2019s relational databases to handle scale","og_description":"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.","og_url":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/","og_site_name":"The GitHub Blog","article_published_time":"2021-09-27T16:02:13+00:00","article_modified_time":"2022-03-02T17:44:58+00:00","og_image":[{"width":1200,"height":630,"url":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","type":"image\/png"}],"author":"Thomas Maurer","twitter_card":"summary_large_image","twitter_image":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","twitter_misc":{"Written by":"Thomas Maurer","Est. reading time":"10 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#article","isPartOf":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/"},"author":{"name":"Thomas Maurer","@id":"https:\/\/github.blog\/#\/schema\/person\/7abf5c5c1cb04b26e5a04536b62c4501"},"headline":"Partitioning GitHub\u2019s relational databases to handle scale","datePublished":"2021-09-27T16:02:13+00:00","dateModified":"2022-03-02T17:44:58+00:00","mainEntityOfPage":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/"},"wordCount":1938,"image":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#primaryimage"},"thumbnailUrl":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","articleSection":["Engineering","Infrastructure"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/","url":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/","name":"Partitioning GitHub\u2019s relational databases to handle scale - The GitHub Blog","isPartOf":{"@id":"https:\/\/github.blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#primaryimage"},"image":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#primaryimage"},"thumbnailUrl":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","datePublished":"2021-09-27T16:02:13+00:00","dateModified":"2022-03-02T17:44:58+00:00","author":{"@id":"https:\/\/github.blog\/#\/schema\/person\/7abf5c5c1cb04b26e5a04536b62c4501"},"description":"In 2019, to meet growth and availability challenges, we set a plan in motion to improve our tooling and ability to partition relational databases.","breadcrumb":{"@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#primaryimage","url":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","contentUrl":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","width":1200,"height":630},{"@type":"BreadcrumbList","@id":"https:\/\/github.blog\/engineering\/infrastructure\/partitioning-githubs-relational-databases-scale\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/github.blog\/"},{"@type":"ListItem","position":2,"name":"Engineering","item":"https:\/\/github.blog\/engineering\/"},{"@type":"ListItem","position":3,"name":"Infrastructure","item":"https:\/\/github.blog\/engineering\/infrastructure\/"},{"@type":"ListItem","position":4,"name":"Partitioning GitHub\u2019s relational databases to handle scale"}]},{"@type":"WebSite","@id":"https:\/\/github.blog\/#website","url":"https:\/\/github.blog\/","name":"The GitHub Blog","description":"Updates, ideas, and inspiration from GitHub to help developers build and design software.","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/github.blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/github.blog\/#\/schema\/person\/7abf5c5c1cb04b26e5a04536b62c4501","name":"Thomas Maurer","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g438a2a1efba5fc2060e7180d2732a58e","url":"https:\/\/secure.gravatar.com\/avatar\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/7c448b5bd1cbb7cb228fe479470021364ef577cb27c1ea6ad1172d52390f67a9?s=96&d=mm&r=g","caption":"Thomas Maurer"},"url":"https:\/\/github.blog\/author\/tma\/"}]}},"jetpack_publicize_connections":[],"jetpack_shortlink":"https:\/\/wp.me\/pamS32-fHJ","jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/github.blog\/wp-content\/uploads\/2021\/04\/Blog_ENGINEERING_for-social.png?fit=1200%2C630","_links":{"self":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/60371","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/users\/1491"}],"replies":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/comments?post=60371"}],"version-history":[{"count":13,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/60371\/revisions"}],"predecessor-version":[{"id":63519,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/posts\/60371\/revisions\/63519"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/media\/57452"}],"wp:attachment":[{"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/media?parent=60371"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/categories?post=60371"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/tags?post=60371"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/github.blog\/wp-json\/wp\/v2\/coauthors?post=60371"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}