Half a Million Working Credentials Are Still Sitting in Public GitHub Repos
Cybersecurity

Half a Million Working Credentials Are Still Sitting in Public GitHub Repos

Truffle Security validated 543,699 live credentials across public GitHub repositories, including one still working sixteen years after it was committed, and the numbers show push protection has not closed the gap enterprises think it closed.

PublishedOctober 2, 2026
Read time5 min read
Share

The scale of the scan is what makes the numbers credible

Truffle Security built its findings on an analysis of The Stack v3, a public-code corpus spanning 58.4 billion files across 224.5 million repositories, not a small sample or a targeted search for a specific credential type. Researchers then took the extra step that separates this research from routine secret-scanning reports: rather than simply flagging strings that look like credentials, they validated each candidate directly against its issuing service on July 27 and 28, 2026, confirming which ones could still actually authenticate.

That validation step is why 543,699 is a meaningfully different number from a typical pattern-matching count of things that merely resemble API keys or tokens. Every one of those credentials worked when tested. Across the full dataset, researchers recorded 1,103,438 separate exposures spanning repositories and individual files, meaning many of the live credentials showed up in more than one place, copied into forks, documentation, or configuration templates committed by different contributors over time.

Push protection is working, just not nearly enough on its own

GitHub's push protection feature, which scans commits for recognizable credential patterns before they reach a public repository, is a genuinely useful control and has clearly blocked a real volume of accidental exposures since its default rollout. But 199,843 of the live credentials Truffle Security found, 36.8 percent of the total, were pushed after that protection was already active by default, meaning more than a third of today's exposures happened on GitHub's watch rather than before the control existed.

Read against the 51.8 percent of live credentials that fell into categories push protection does not cover particularly well, the picture is a control that catches an important slice of the problem rather than one that has solved it. Enterprises that adopted push protection and treated secrets exposure as a closed risk from that point forward are working from an outdated threat model, and this data is a direct argument for layering continuous repository scanning on top of point-in-time commit checks.

Sixteen years is not an outlier worth dismissing

The single most striking data point in the report is the oldest verified credential: tied to a file last modified in June 2009, it was still operational more than sixteen years later when Truffle Security tested it in July 2026. A credential surviving that long, through at least one and probably several changes of engineering leadership, multiple security program overhauls, and likely more than one compliance audit, is a direct measure of how rarely organizations actually rotate secrets once they are issued.

The median exposure duration across the full dataset, 784 days, tells the broader version of the same story. A credential sitting exposed and functional for more than two years on average is not primarily a discovery problem, security teams generally have tools capable of finding these if pointed at the right repositories. It is a rotation and lifecycle management problem, where issuing a credential is treated as a one-time setup task rather than an asset with an ongoing maintenance obligation.

Why database strings survive and tokens mostly don't

The survival-rate data by credential type is the most actionable part of the whole report. Postgres connection strings survived at an 88 percent rate and MySQL strings at 75 percent, while npm tokens survived at roughly 0.001 percent and GitHub tokens at 0.36 percent. That gap is not random, it reflects which credential types have mature, well-adopted rotation tooling built into their ecosystems and which do not.

npm and GitHub both invested heavily in automated token revocation and short-lived credential patterns, and the survival data shows that investment paying off directly in practice. Database connection strings, by contrast, are frequently hand-rotated, embedded in configuration files that get copied between environments, and treated as infrastructure rather than as identity, which is exactly the category of thinking that needs to change given how long these specific credentials evidently survive once exposed.

The specific exposure categories enterprises should check first

Beyond the headline numbers, the report breaks out specific high-volume categories worth checking directly against your own organization's repositories and forks: 69,041 live Google Cloud service account credentials, 51,067 live MongoDB connection strings, and 33,343 live Google API keys. Any organization using these services should treat a search of its own public and recently-public repository history as a same-week priority rather than a future audit item, given how directly this data maps to active, exploitable access.

That search needs to cover forks and archived repositories as well as active ones. A credential committed years ago to a repository that has since been archived, made private, or forked by a departed contributor does not stop working just because the repository fell out of active use, and Truffle Security's own methodology, scanning a broad historical corpus rather than only current public repositories, is the reason this research surfaced so much that routine, present-day-only scanning tools would have missed.

What this changes about your secrets management program

The practical takeaway is that secrets management needs to be measured by rotation cadence and exposure duration, not by whether a scanning tool is deployed. A program that can answer how old its oldest active, unrotated credential is, and can point to an automated process that would catch a 784-day-old exposure well before that number got anywhere close, is in meaningfully better shape than one that simply confirms a secret-scanning product is installed and configured somewhere in the pipeline.

Prioritize the credential types this data shows survive longest, database connection strings and cloud service account credentials in particular, for short-lived or automatically rotated alternatives wherever your infrastructure supports them. The tooling already exists for most major cloud providers. The gap this report documents is adoption and operational discipline, not a missing technical capability, which makes it a solvable problem rather than a research frontier.

Tagged#news#security#cybersecurity#breach#cisa#ransomware#zero-day#supply-chain#ai-security#github#secrets-management#credential-exposure#push-protection#truffle-security#developer-security