Link Analysis

Anand Rajaraman; Jeffrey David Ullman

doi:10.1017/CBO9781139058452.006

This chapter is part of a book that is no longer available to purchase from Cambridge Core

5 - Link Analysis

Anand Rajaraman and

Jeffrey David Ullman

Show author details

Anand Rajaraman: Affiliation:
WalmartLabs
Jeffrey David Ullman: Affiliation:
Stanford University, California

Book contents

Get access

Summary

One of the biggest changes in our lives in the decade following the turn of the century was the availability of efficient and accurate Web search, through search engines such as Google. While Google was not the first search engine, it was the first able to defeat the spammers who had made search almost useless. Moreover, the innovation provided by Google was a nontrivial technological advance, called “PageRank.” We shall begin the chapter by explaining what PageRank is and how it is computed efficiently.

Yet the war between those who want to make the Web useful and those who would exploit it for their own purposes is never over. When PageRank was established as an essential technique for a search engine, spammers invented ways to manipulate the PageRank of a Web page, often called link spam. That development led to the response of TrustRank and other techniques for preventing spammers from attacking PageRank. We shall discuss TrustRank and other approaches to detecting link spam.

Finally, this chapter also covers some variations on PageRank. These techniques include topic-sensitive PageRank (which can also be adapted for combating link spam) and the HITS, or “hubs and authorities” approach to evaluating pages on the Web.

PageRank

We begin with a portion of the history of search engines, in order to motivate the definition of PageRank, a tool for evaluating the importance of Web pages in a way that it is not easy to fool. We introduce the idea of “random surfers,” to explain why PageRank is effective.

Type: Chapter
Information: Mining of Massive Datasets , pp. 139 - 175

DOI: https://doi.org/10.1017/CBO9781139058452.006 [Opens in a new window]

Publisher: Cambridge University Press

Print publication year: 2011

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Book contents

5 - Link Analysis

Summary

Access options

Save book to Kindle

Save book to Dropbox

Save book to Google Drive