How GitHub Uses Elasticsearch to Bring Semantic Search to 395 Million Code Repositories
GitHub, the world’s largest code host serving 180 million developers, deployed Elasticsearch on Elastic Cloud to add semantic search across more than 395 million repositories and billions of documents. The system handles natural-language queries from both human developers and AI agents, dramatically reducing zero-hit search results and improving click-through rates. A team of five to six engineers runs the entire search platform at that scale, with BBQ vector compression reducing infrastructure costs 32x.
Tools & Technologies
1AI Categories
Challenge
GitHub’s keyword-based search failed to handle the natural-language queries developers increasingly use, and broke entirely for AI agents and assistants that interact with GitHub data as first-class clients, leaving users with zero-hit results.
Solution
Elasticsearch on Elastic Cloud was deployed to power semantic search across billions of documents, using vector embeddings and BBQ compression to handle natural-language queries from humans and AI systems at scale, with Kibana enabling the engineering team to iterate quickly.
Full Story
GitHub is where the world builds software. More than 180 million developers at 4 million organizations — including 90% of the Fortune 100 — rely on it to create, store, and share code. That means GitHub manages more than 395 million repositories and billions of documents covering source code, patch notes, discussions, and wikis.
Access 455+ AI use cases, 427+ tools, and adoption signal rankings.