Google Leaks: Google Document Leak Comes to Light

Marketing Digital ·

SEO Agency Uruguay - Web Positioning Agency Uruguay

Who said the world of SEO was boring? Learn all the details of the latest Google algorithm data leak and how it impacts organic web positioning.

What happened

Despite high security standards, data leaks at Google are nothing new. However, so far in the 21st century, few have been as significant as the one that came to light on Monday, May 27.

In late March of this year, documents from the Google Search API content warehouse, originally intended for internal use, were leaked from GitHub. These documents migrated from GitHub to the Hexdocs repository and circulated through other sources until they finally became public.

On May 5, Rand Fishkin received an email from an anonymous source who later identified themselves as Efran Azimi. In his email, Azimi claimed to have access to internal documents from Google's search service.

Although the leak of the Google repository was reportedly fixed on May 7, the documents remained on publicly accessible pages. Therefore, on May 24, Fishkin and Azimi had a video call, during which Azimi showed the documentation and explained his motivations to Fishkin.

How did it all continue? Fishkin, a self-proclaimed retired SEO who in the early 2010s was the king of SEO at MOZ with his “whiteboard friday” that enchanted all of us wannabe SEOs who admired his style and knowledge.

Rand Fishkin is now the Founder and CEO of the audience analysis platform Sparktoro, and confirmed with former Google employees that these were legitimate documents.

His next step was to turn to Mike King (a technical SEO expert, Founder & CEO of iPullRank) to decipher the documentation. Once King completed an initial analysis of the document and shared the results with Fishkin, both published articles about the topic… And the rest is history.

What Google documentation was leaked?

Although the documents are dated August 2023, it is reasonable to believe they were still valid as of March of this year. They constitute a kind of “guide” for the Xooglers (Google employees) team in the search department.

Through codes and brief descriptions, the documentation explains the attributes and API modules of Google Search.

That is, the interface that mediates between two systems so they share information and functionalities.

The information covers 2,596 modules of the Google Search API. These are specified in types, functions, and 14,014 attributes (i.e., characteristics for ranking) that the interface considers when collecting data.

Thus, with these documents, it is possible to know what data Google retrieves from its users' searches and their behavior before, during, and after these searches.

Apparently, information is recorded about the links and content of a page, and about user interactions with them. It is based on this data and Google's algorithms that a page will rank better or worse in the SERPs.

Thus, it seems this leak is a gift from heaven for SEO specialists.

However, many of the entries refer to internal Google pages, which can only be accessed with a company credential. This means that King's analysis and the conclusions he and Fishkin drew should not be taken as definitive or total.

Both experts, in fact, acknowledged a series of limitations in their work with the documentation. First, although the leaked documents indicate which data they record, they do not distinguish all the elements that influence a page's ranking.

It is known, however, that some are obsolete for organizing the SERPs. On the other hand, the documentation does not indisputably identify what information is used nor in what ways. It is also not possible to determine which characteristics have more weight than others when ranking a web page.

Even so, it should be noted that to analyze the documents, they drew on their extensive experience as SEOs and the practices of other professionals in the field.

They also considered ranking systems similar to those explained by Google. In short: without being able to make precise claims, they reached more than likely conclusions, perfect for getting a clearer idea of how to apply SEO to achieve ideal results.

The heart of the matter

Now, what is serious about this matter for those of us who practice SEO? Or rather, why are the leaked data so controversial? On one hand, because more than once Google executives denied the influence of certain metrics on web ranking.

Likewise, they harshly criticized SEO professionals (Rand Fishkin among them) who claimed that Google organized the SERPs based on Chrome data, the time a user spends on a page, sandboxes, among other things.

On the other hand, the documents reveal that not all SEO practices are really useful when ranking a page. The documentation presents a ranking system where writing for search engines matters less than addressing users… Unless, for example, one is backed by a recognized brand, which will attract clicks like flies to honey.

In other words: regarding the page ranking, Google did not share even half of what happens in its search engine system.

Moreover, it wasn't just the lack of information that was problematic, but also the way Google handled the situation.

Rather than not sharing even half, it led us in circles, sending novice SEOs down wrong paths and confusing them in the great labyrinth of the cyber minotaur.

This complexity and lack of transparency made the task of optimizing pages for the search engine even more difficult, leaving many professionals frustrated and disoriented.

Leaked documents: Key points for SEO

Now let's get to what should be your biggest concern: from everything that was leaked, what does an SEO expert need to know? While we can't know it all yet, in their articles Fishkin and King (in a very technical way) recovered several aspects of the Google API that affect SEO. Some of the most important ones are:

  • Navboost. It is one of the search engine's internal algorithms. Together with Glue (another algorithm), they order the SERP entries according to the quantity and types of clicks they receive, over a range of up to 13 months. The more and better clicks a page receives, the higher it ranks, as it signals to the search engine that the link is trustworthy.
  • Chrome data. Google Search collects information about user behavior in Chrome. It uses this data, among other things, to define the sitelinks it will present in the SERPs. Which URLs will appear? Those with which users interact the most. That is, where they click more and stay longer.
  • Quality raters. The Google Search documentation confirms that the search engine uses information provided by its quality raters through the Ewok platform to determine search results.
  • Sandbox. Even though Google denied it, it had been suspected for some time that the search engine used “sandboxes” to determine the trustworthiness of new domains. Thus, no matter how good the SEO of a page is, while it goes through this algorithm, it will be as if it didn't exist in searches. It's a matter of patience, at least supposedly.
  • Twiddlers. They are “functions” of re-ranking that are applied after Google Search's main algorithms. They allow, in this way, specifying the entries that will be shown in the SERPs.

The above points explain some of the tools that Google Search uses to filter the pages it will rank in the SERPs. But King and Fishkin do not stop there: they also mention certain “internal” aspects that affect page ranking. These are:

Brand popularity

According to what the leaked documentation shows, the best SEO has little to do against an established brand where users will inevitably click. Also, it was discovered that Google Search identifies pages that are a small personal site. Given its tendency to prioritize large companies, this is not very encouraging for small-scale businesses.

Titles

When analyzing the documents, King found a mention of a titlematchScore. For SEO purposes, this means that when generating SERPs, Google will look for the title of the pages and the search query to match.

Bold and font size

According to King's deciphering, through the attributes avgTermWeight and fontsize, Google Search registers bold terms or those with larger font sizes. Consequently, it is logical to assume that these typographic modifications can help achieve a good level of web positioning.

Dates

The leaked documents include the attributes bylineDate, syntacticDate, and semanticDate. Respectively, these compare the dates appearing in the page, extracted from the title or URL, and from the content of the page. From this, King infers that for a page's ranking in the SERPs, it is important that the different dates it includes match.

Finally, some factors that negatively influence a page's ranking in search results are:

  • Potential user dissatisfaction, determined by the page receiving a low number of clicks.
  • If the page is “global” or not associated with a specific location.
  • UX or navigation issues.
  • Inclusion of pornographic elements.
  • Using exact match domains (e.g., “www.zapatos-para-mujer.com”).

What can we say?: The user is key

Being objective and broadly speaking, the most serious thing about this matter is that the leaked documentation openly contradicts many public statements made by Google executives.

To make matters worse, with some of those statements, those professionals who sought to disprove certain assumptions about Google's internal mechanisms were discredited. Of course, as he who laughs last laughs best, this data leak took care of clarifying who was telling (or deducing) the truth.

On the other hand, the seriousness of the matter lies in the deception suffered by those of us who dedicate ourselves to SEO. Certainly, the mystery of many of the measures applied by the search engine serves to disorient spammers and ensure the quality of pages.

However, the leaked documents show that sometimes the weight of a brand can leave quality in the background. They also show that ranking is not as “organic” as one might suppose. And that it depends less on the skill of an SEO professional than on what Google considers many or few clicks.

In any case, the leaked information serves to confirm a trend that Google and other browsers have been assuming in recent years: produce for the user, not for a robot. Just look at the elements that contribute to ranking in the SERPs. User experience and navigability, fonts that facilitate reading, quality content that matches the search performed…

The focus is on what the user needs. And satisfactorily meeting that need is the most organic measure one can take to rank well today. After all, quality clicks are not given by Google (or maybe they are, but we can't confirm it), but by those who browse the web.

None of this is new, nor does it imply a very big change for those who were already applying SEO practices oriented to user experience in their production. In summary: the leak from the Google Search API content warehouse confirmed (among other things) that the consumer is key. But we already knew that.

How to improve your SEO practices for Google?

Now, adapting to this new paradigm overnight can be a bit difficult. Especially if you don't work with marketing specialists with a solid team of SEO experts who can advise you on the best ways to reach the user.

At a general level, for now, there are some issues that (beyond what has already been mentioned) help generate optimal ranking in the SERPs:

  • First, it is essential to strengthen organic positioning and establish your brand using resources that go beyond SEO. Google Ads, newsletter campaigns, social media presence, word of mouth… No matter which medium you choose, the important thing is that it produces specific searches for your page, leading to quality clicks.
  • Second, according to Mike King, Google algorithms tend to consider “fresh” content as quality. Therefore, one way to improve the chances of achieving good organic positioning is to keep your pages updated. Additionally, when including external links, the ideal is that they too are recent.
  • A third way to improve your SEO practices is design your pages with the user in mind. Identify the paths they will take to reach it and once they have entered; the information they will look for; the actions they will carry out… Clearly, this is easier said than done. That's why the fourth point to consider can only be one:
  • Are you lost? Can't quite figure out how to proceed? Then don't be afraid to consult an agency of specialists in web positioning, like Evox.

We have a broad team of experts in SEO, SEM, web design, social media, and more. We offer unique solutions, tailored to your needs. Over ten years of experience in the field backs us to help you achieve the search engine presence your projects need. Don't let Google's algorithms swallow your projects. Come to the surface with a couple of clicks: contact us.