Uses of Class
org.apache.nutch.protocol.Content

Packages that use Content
Package
Description
Text document language identifier.
Crawl control code and tools to run the crawler.
A microformats Rel-Tag Parser/Indexer/Querier plugin.
The Parse interface and related classes.
Parse wrapper to run external command to do the parsing.
Parse RSS feeds.
Parse filter to extract headings (h1, h2, etc.) from DOM parse tree.
An HTML document parsing plugin.
Parser and parse filter plugin to extract all (possible) links from JavaScript files and embedded JavaScript code snippets.
Parse filter to extract meta tags: keywords, description, etc.
Parse various document formats with help of Apache Tika.
Parse ZIP files: embedded files are recursively passed to appropriate parsers.
Adds serialized DOM to parse data, useful for debugging, to understand how the parser implementation interprets a document (not only HTML).
Html Parse filter that classifies the outlinks from the parseresult as relevant or irrelevant based on the parseText's relevancy (using a training file where you can give positive and negative example texts see the description of parsefilter.naivebayes.trainfile) and if found irrelevent it gives the link a second chance if it contains any of the words from the list given in parsefilter.naivebayes.wordlist.
RegexParseFilter.
Classes related to the Protocol interface, see also org.apache.nutch.net.protocols.
Protocol plugin which supports retrieving local file resources.
Protocol plugin which supports retrieving documents via the ftp protocol.
Common API used by HTTP plugins (http, httpclient, etc.)
The ScoringFilter interface.
Scoring filter to stop crawling at a configurable depth (number of "hops" from seed URLs).
Scoring filter used in conjunction with WebGraph.
Metadata Scoring Plugin
Scoring filter implementing a variant of the Online Page Importance Computation (OPIC) algorithm.
 
Implements the cosine similarity metric for scoring relevant documents
URL Meta Tag Scoring Plugin
A segment stores all data from on generate/fetch/update cycle: fetch list, protocol status, raw content, parsed content, and extracted outgoing links.
Miscellaneous tools.
Miscellaneous utility classes.
Sample plugins that parse and index Creative Commons metadata.