How Pushing Bad Data Google S Latest Black Eye

Google stopped counting, or at least showing publicly, the number of pages it pointed to in September 05, after a “yardstick competition” with the school yard with a rival Yahoo. That number reached nearly 8 billion pages before it was removed from the homepage. News has recently exploded in various SEO forums that Google suddenly, over the past few weeks, added another few billion pages to the index. This may sound like a reason to celebrate, but this "accomplishment" would not be a good indication of a successful search engine.

 

What was buzzing with people was the kind of few billion new pages. It was crude spam - which contained Pay-Per-Click (PPC) ads, discarded content, and, in many cases, was clearly visible in search results. Bring out the oldest and most established sites in the process. A Google representative has responded to the forums in the matter by calling it “bad data compression,” something that has been the subject of various complaints in the SEO community.

 

How can anyone trick Google into identifying so many spam pages in such a short period of time? I will give you an idea of ​​the high level of process, but do not be too happy. Just like a nuclear bomb will not teach you how to do something real, you will not be able to run and do it yourself after reading this article. Yet, it does make an interesting story, showing the serious problems that are found in growing up in the world-famous search engine.

 

Dark Night and Storm

 

Our story begins in the heart of Moldva, located between Romania and Ukraine. Amidst the ban on local vampire attacks, the wise local man had a clever idea and ran with it, almost away from the vampires… His vision was to exploit the way Google treats subdomains, and not just a little, but in a big way.

 

At the heart of the problem is that at the moment, Google treats subdomains in the same way as it treats full-fledged domains - like separate entities. This means it will add a subdomain homepage to the index and come back at some point to make a "deep crawl." Deep crawling is just a spider following links from the homepage to the depth of the site until we find everything or stop and come back later to find out more.

 

In short, subdomain is a “third-tier domain.” You may have seen them before, they look like this: subdomain.domain.com. Wikipedia, for example, uses them for languages; the English version says “en.wikipedia.org”, the Dutch version says “nl.wikipedia.org.” Subdomains are one way to organize large sites, unlike many texts or even completely different domain names.

 

So, we have a kind of page that Google will point to, almost "no questions asked." Surprisingly, no one used this situation immediately. Some analysts believe that the reason for this may be “bad” introduced after the recent “Big Daddy” update. Our Eastern European friend has integrated some servers, content scrapers, spambots, PPC accounts, and the most important, most inspired documents, and put them together like this ...

 

Five Million Used - And Counting ...

 

First, our hero here created scripts for his servers that when GoogleBot passed, began to produce an infinite number of subdomains, all with a single page containing rich keywords, keyword links, and PPC ads for those keywords. Spambots are being sent to add GoogleBot smells by transmitting and commenting on spam to tens of thousands of blogs worldwide. Spambots offer a wide range of settings, and it does not take much to make the dominos fall.

 

GoogleBot detects spam links and, as its purpose for life, follows them to the network. Once GoogleBot is sent to the web, servers running servers simply continue to produce page-by-page, all with a different subdomain, all with keywords, scratched content, and PPC ads. These pages are listed, and you have suddenly acquired Google's index of 3-5 billion pages that weigh less than 3 weeks.

 

Reports indicate, initially, PPC ads on these pages were from Adsense, Google's PPC service. Surprisingly, Google benefits financially from all the expense charged by Adsense users as it emerges from these billions of spam pages. Adsense revenue from this effort was a point, after all. Cram on so many pages that, with the power of numbers, people can find and click ads on those pages, making the spam sender a very good profit in a very short time.

 

Billions or Millions? What Is Broken?

 

The word of this success spread like wildfire from DigitalPoint sites. It is spreading like wildfire in the SEO community, to clarify. The "public", meanwhile, is not in the loop, and will likely remain so. The Google developer response appeared in the Threadwatch series on the topic, calling it "bad data compression". In fact, the company's lineup was that they had not added 5 billion pages. Recent claims include guarantees that the issue will be resolved algorithmically. Those who follow the status quo (by tracking the known spam domains they are using) only see that Google is automatically indexing them.

 

Tracking is done using the "site:" command. The command that, in theory, shows the total number of index pages on the site you specify after the colon. Google has already acknowledged that there are problems with this command, and "5 billion pages", it seems, is looking for it, it's just a symbol of it. These problems extend beyond the site only: the command, but the display of a number of results for many queries, some of which feel as if they are not very accurate and in some cases very volatile. Google agrees to identify some of these sub-spam domains, but so far they have not provided additional numbers to counter the 3-5 billion previously displayed on the site: the order.

Enjoyed this article? Stay informed by joining our newsletter!

Comments

You must be logged in to post a comment.

About Author