Topic 4 W4 - Text Processing

Uploaded by

VISALINI VIJAYAN

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

19 views42 pages

Topic 4 W4 - Text Processing

Uploaded by

VISALINI VIJAYAN

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPTX, PDF, TXT or read online on Scribd

You are on page 1/ 42

Search Engines

Information Retrieval in Practice

All slides ©Addison Wesley, 2008

Processing Text
• Converting documents to index terms
• Why?
– Matching the exact string of characters typed by
the user is too restrictive
• i.e., it doesn’t work very well in terms of effectiveness
– Not all words are of equal value in a search
– Sometimes not clear where words begin and end
• Not even clear what a word is in some languages
– e.g., Chinese, Korean
Text Statistics
• Huge variety of words used in text but
• Many statistical characteristics of word
occurrences are predictable
– e.g., distribution of word counts
• Retrieval models and ranking algorithms
depend heavily on statistical properties of
words
– e.g., important words occur often in documents
but are not high frequency in collection
Zipf’s Law
• Distribution of word frequencies is very
skewed
– a few words occur very often, many words hardly
ever occur
– e.g., two most common words (“the”, “of”) make
up about 10% of all word occurrences in text
documents
Zipf’s Law
• the frequency of any word is inversely proportional to its
rank in the frequency table.
• most frequent word will occur approximately twice as often
as the second most frequent word,
• three times as often as the third most frequent word, etc.
• For example in a doc., the word "the" is the most frequently
occurring word, and by itself accounts for nearly 7% of all
word occurrences (69,971 out of slightly over 1 million).
• True to Zipf's Law, the second-place word "of" accounts for
slightly over 3.5% of words (36,411 occurrences), followed
by "and" (28,852).
Zipf’s Law
Vocabulary Growth
• As corpus grows, so does vocabulary size
– Fewer new words when corpus is already large
• Observed relationship (Heaps’ Law):

v = k.nβ
where v is vocabulary size (number of unique words),
n is the number of words in corpus, k, β are
parameters that vary for each corpus (typical values
given are 10 ≤ k ≤ 100 and β ≈ 0.5)
AP89 Example

k β
Heaps’ Law Predictions
• Predictions for TREC collections are accurate
for large numbers of words
– e.g., first 10,879,522 words of the AP89 collection
scanned
– prediction is 100,151 unique words
– actual number is 100,024
• Predictions for small numbers of words (i.e.
< 1000) are much worse
GOV2 (Web) Example
Web Example
• Heaps’ Law works with very large corpora
– new words occurring even after seeing 30 million!
– parameter values different than typical TREC
values
• New words come from a variety of sources
• spelling errors, invented words (e.g. product, company
names), code, other languages, email addresses, etc.
• Search engines must deal with these large and
growing vocabularies
Estimating Result Set Size

• How many pages contain all of the query terms?

• For the query “a b c”:
fabc = N · fa/N · fb/N · fc/N = (fa · fb · fc)/N2

• Assuming that terms occur independently

• fabc is the estimated size of the result set
• fa, fb, fc are the number of documents that terms a, b, and
c occur in
• N is the number of documents in the collection
GOV2 Example

(fa · fb )/N

Collection size (N) is 25,205,179

Tokenizing
• Forming words from sequence of characters
• Surprisingly complex in English, can be harder
in other languages
• Early IR systems:
– any sequence of alphanumeric characters of
length 3 or more
– terminated by a space or other special character
– upper-case changed to lower-case
Tokenizing
• Example:
– “Bigcorp's 2007 bi-annual report showed profits
rose 10%.” becomes
– “bigcorp 2007 annual report showed profits rose”
• Why? Too much information lost
– Small decisions in tokenizing can have major
impact on effectiveness of some queries
Tokenizing Problems
• Small words can be important in some queries,
usually in combinations
• xp, ma, pm, ben e king, el paso, master p, gm, j lo, world
war II
• Both hyphenated and non-hyphenated forms of
many words are common
– Sometimes hyphen is not needed
• e-bay, wal-mart, active-x, cd-rom, t-shirts
– At other times, hyphens should be considered either
as part of the word or a word separator
• winston-salem, mazda rx-7, e-cards, pre-diabetes, t-mobile,
spanish-speaking
Tokenizing Problems
• Special characters are an important part of tags,
URLs, code in documents
• Capitalized words can have different meaning
from lower case words
– Bush, Apple
• Apostrophes can be a part of a word, a part of a
possessive, or just a mistake
– rosie o'donnell, can't, don't, 80's, 1890's, men's straw
hats, master's degree, england's ten largest cities,
shriner's
Tokenizing Problems
• Numbers can be important, including decimals
– nokia 3250, top 10 courses, united 93, quicktime
6.5 pro, 92.3 the beat, 288358
• Periods can occur in numbers, abbreviations,
URLs, ends of sentences, and other situations
– I.B.M., Ph.D., cs.umass.edu, F.E.A.R.
• Note: tokenizing steps for queries must be
identical to steps for documents
Tokenizing Process
• First step is to use parser to identify
appropriate parts of document to tokenize
• Defer complex decisions to other components
– word is any sequence of alphanumeric characters,
terminated by a space or special character, with
everything converted to lower-case
– everything indexed
– example: 92.3 → 92 3 but search finds documents
with 92 and 3 adjacent
Tokenizing Process
• Not that different than simple tokenizing
process used in past
• Examples of rules used with TREC
– Apostrophes in words ignored
• o’connor → oconnor bob’s → bobs
– Periods in abbreviations ignored
• I.B.M. → ibm Ph.D. → ph d
Stopping
• Function words (determiners, prepositions)
have little meaning on their own
• High occurrence frequencies
• Treated as stopwords (i.e. removed)
– reduce index space, improve response time,
improve effectiveness
• Can be important in combinations
– e.g., “to be or not to be”
Stopping
• Stopword list can be created from high-
frequency words or based on a standard list
• Lists are customized for applications, domains,
and even parts of documents
– e.g., “click” is a good stopword for anchor text
• Best policy is to index all words in documents,
make decisions about which words to use at
query time
Stemming
• Many morphological variations of words
– inflectional (plurals, tenses)
– derivational (making verbs nouns etc.)
• In most cases, these have the same or very
similar meanings
• Stemmers attempt to reduce morphological
variations of words to a common stem
– usually involves removing suffixes
• Can be done at indexing time or as part of
query processing (like stopwords)
Stemming
• Generally a small but significant effectiveness
improvement
– can be crucial for some languages
– e.g., 5-10% improvement for English, up to 50% in
Arabic

Words with the Arabic root ktb

Stemming
• Two basic types
– Dictionary-based: uses lists of related words
– Algorithmic: uses program to determine related
words
• Algorithmic stemmers
– suffix-s: remove ‘s’ endings assuming plural
• e.g., cats → cat, lakes → lake, wiis → wii
• Many false negatives: supplies → supplie
• Some false positives: ups → up
Porter Stemmer
• Algorithmic stemmer used in IR experiments
since the 70s
• Consists of a series of rules designed to the
longest possible suffix at each step
• Effective in TREC
• Produces stems not words
• Makes a number of errors and difficult to
modify
Krovetz Stemmer
• Hybrid algorithmic-dictionary
– Word checked in dictionary
• If present, either left alone or replaced with “exception”
• If not present, word is checked for suffixes that could be
removed
• After removal, dictionary is checked again
• Produces words not stems
• Comparable effectiveness
• Lower false positive rate, somewhat higher false
negative
Stemmer Comparison
Phrases
• Many queries are 2-3 word phrases
• Phrases are
– More precise than single words
• e.g., documents containing “black sea” vs. two words
“black” and “sea”
– Less ambiguous
• e.g., “big apple” vs. “apple”
• Can be difficult for ranking
• e.g., Given query “fishing supplies”, how do we score
documents with
– exact phrase many times, exact phrase just once, individual words
in same sentence, same paragraph, whole document, variations on
words?
Document Structure and Markup
• Some parts of documents are more important
than others
• Document parser recognizes structure using
markup, such as HTML tags
– Headers, anchor text, bolded text all likely to be
important
– Metadata can also be important
– Links used for link analysis
Example Web Page
hypertext
Example Web Page

hypertext
Link Analysis
• Links are a key component of the Web
• Important for navigation, but also for search
– e.g., <a href="http://example.com" >Example
website</a>
– “Example website” is the anchor text
– “http://example.com” is the destination link
– both are used by search engines
Anchor Text
• Used as a description of the content of the
destination page
– i.e., collection of anchor text in all links pointing to
a page used as an additional text field
• Anchor text tends to be short, descriptive, and
similar to query text
• Retrieval experiments have shown that anchor
text has significant impact on effectiveness for
some types of queries
PageRank
• Billions of web pages, some more informative
than others
• Links can be viewed as information about the
popularity (authority?) of a web page
– can be used by ranking algorithm
• Inlink count could be used as simple measure
• Link analysis algorithms like PageRank provide
more reliable ratings
Dangling Links
• Random jump prevents getting stuck on
pages that
– do not have links
– contains only links that no longer point to
other pages
– have links forming a loop
• Links that point to the first two types of
pages are called dangling links
– may also be links to pages that have not yet
been crawled
Link Quality
• Link quality is affected by spam and other
factors
– e.g., link farms to increase PageRank
– trackback links in blogs can create loops
– links from comments section of popular blogs
• Blog services modify comment links to contain
rel=nofollow attribute
• e.g., “Come visit my <a rel=nofollow
href="http://www.page.com">web page</a>.”
Trackback Links
Internationalization
• 2/3 of the Web is in English
• About 50% of Web users do not use English as
their primary language
• Many (maybe most) search applications have
to deal with multiple languages
– monolingual search: search in one language, but
with many possible languages
– cross-language search: search in multiple
languages at the same time
Internationalization
• Many aspects of search engines are language-
neutral
• Major differences:
– Text encoding (converting to Unicode)
– Tokenizing (many languages have no word
separators)
– Stemming
• Cultural differences may also impact interface
design and features provided
Chinese “Tokenizing”
END

StotraNidhi Telugu 15-Books Combo
No ratings yet
StotraNidhi Telugu 15-Books Combo
1 page
Reset Blu Ray Samsung BD-F5100
0% (1)
Reset Blu Ray Samsung BD-F5100
5 pages
Chap 4
No ratings yet
Chap 4
76 pages
Chapter 4
No ratings yet
Chapter 4
72 pages
Chapter 4 - Processing Text
No ratings yet
Chapter 4 - Processing Text
7 pages
IR Chapter 2 Text Operations
No ratings yet
IR Chapter 2 Text Operations
25 pages
Chapter-2 - Automatic Text Anlysis
No ratings yet
Chapter-2 - Automatic Text Anlysis
67 pages
Text Operations 2021
No ratings yet
Text Operations 2021
45 pages
Chapter 2 (Information Storage & Retrieval)
No ratings yet
Chapter 2 (Information Storage & Retrieval)
56 pages
Lecture 3
No ratings yet
Lecture 3
70 pages
2T-Inverted Index
No ratings yet
2T-Inverted Index
54 pages
2 - Text Operation - 1
No ratings yet
2 - Text Operation - 1
28 pages
Tokenization: Token Normalization Is The Process of Canonicalizing Tokens So That Matches Occur
No ratings yet
Tokenization: Token Normalization Is The Process of Canonicalizing Tokens So That Matches Occur
3 pages
2-Text Operations - New
No ratings yet
2-Text Operations - New
39 pages
2 - Text Operation
No ratings yet
2 - Text Operation
35 pages
Chapter 2 Text Operations
No ratings yet
Chapter 2 Text Operations
37 pages
2 - Text Operation
No ratings yet
2 - Text Operation
45 pages
Chapter Two IR
No ratings yet
Chapter Two IR
44 pages
2&3 Text Operation
No ratings yet
2&3 Text Operation
65 pages
CH 2 - Text Operation
No ratings yet
CH 2 - Text Operation
38 pages
Multimedia Information Retrieval (CSC 545) : The Problem of IR
No ratings yet
Multimedia Information Retrieval (CSC 545) : The Problem of IR
29 pages
2 TextOperations
No ratings yet
2 TextOperations
54 pages
Chap 6
No ratings yet
Chap 6
70 pages
Bulu
No ratings yet
Bulu
47 pages
1-Getting Started With ELK
No ratings yet
1-Getting Started With ELK
44 pages
Completed UNIT-III 20.9.17
No ratings yet
Completed UNIT-III 20.9.17
61 pages
Midterm 1
No ratings yet
Midterm 1
5 pages
IRS Chapter 2
No ratings yet
IRS Chapter 2
57 pages
My M-7
No ratings yet
My M-7
44 pages
2 Text Operations
No ratings yet
2 Text Operations
32 pages
MSC IR 2021
100% (1)
MSC IR 2021
188 pages
Lecture 3-Term Vocabulary and Posting Lists
No ratings yet
Lecture 3-Term Vocabulary and Posting Lists
38 pages
Module 5 - Information Retrieval and Lexical Resources
0% (1)
Module 5 - Information Retrieval and Lexical Resources
80 pages
2 - Text Operation
No ratings yet
2 - Text Operation
47 pages
Lecture 3-Term Vocabulary and Posting Lists
No ratings yet
Lecture 3-Term Vocabulary and Posting Lists
26 pages
Indexing Processes (Text Transformation)
No ratings yet
Indexing Processes (Text Transformation)
10 pages
2 Text Operation
No ratings yet
2 Text Operation
46 pages
IR Chapter 2
No ratings yet
IR Chapter 2
37 pages
Chapter 2 Part 1 & 2
No ratings yet
Chapter 2 Part 1 & 2
58 pages
Text-Processing
No ratings yet
Text-Processing
70 pages
2 - Text Operations
No ratings yet
2 - Text Operations
56 pages
Chapter 4 IR
No ratings yet
Chapter 4 IR
56 pages
Internet Searching: Crawling Is Conceptually Quite Simple: Starting at Some Well-Known Sites On The Web
No ratings yet
Internet Searching: Crawling Is Conceptually Quite Simple: Starting at Some Well-Known Sites On The Web
4 pages
02 Text Operation
No ratings yet
02 Text Operation
52 pages
Lecture 2 - Web Search
No ratings yet
Lecture 2 - Web Search
40 pages
Unit 3 - Basic Tokenizing, Indexing, and Implementation of Vector-Space Retrieval
No ratings yet
Unit 3 - Basic Tokenizing, Indexing, and Implementation of Vector-Space Retrieval
8 pages
17 Assignment 4 RSS PDF
No ratings yet
17 Assignment 4 RSS PDF
4 pages
Assignment 1
No ratings yet
Assignment 1
23 pages
Chapter - 6 Part 1
No ratings yet
Chapter - 6 Part 1
21 pages
2-Google Search Basics
No ratings yet
2-Google Search Basics
4 pages
ch2 - Text Operations and Automatic Indexing
No ratings yet
ch2 - Text Operations and Automatic Indexing
20 pages
Lecture 3 - Terms, Postings, Dictionaries, and Tolerant Retrieval
No ratings yet
Lecture 3 - Terms, Postings, Dictionaries, and Tolerant Retrieval
77 pages
CSE 435/535 Information Retrieval: Chapter 2: Tokenization, Stemming, Lemmatization
No ratings yet
CSE 435/535 Information Retrieval: Chapter 2: Tokenization, Stemming, Lemmatization
48 pages
IR Problem: Introduction To Information Retrieval Outline
No ratings yet
IR Problem: Introduction To Information Retrieval Outline
11 pages
CL - Lec 6
No ratings yet
CL - Lec 6
28 pages
Chapter 3 IR
No ratings yet
Chapter 3 IR
56 pages
Faculty Name: Dr. Humera Khanam Subject Name:NLP
No ratings yet
Faculty Name: Dr. Humera Khanam Subject Name:NLP
206 pages
Chapter 3,4, 5 and 6
No ratings yet
Chapter 3,4, 5 and 6
145 pages
AI6122 Topic 3.1 - Index
No ratings yet
AI6122 Topic 3.1 - Index
40 pages
Chapter Two - Text Operations and Automatic Indexing: 2.1. Text Acquisition Via Crawler
No ratings yet
Chapter Two - Text Operations and Automatic Indexing: 2.1. Text Acquisition Via Crawler
19 pages
Schematron: A language for validating XML
From Everand
Schematron: A language for validating XML
Erik Siegel
No ratings yet
Key & Common Swedish Words A Vocabulary List of High Frequency Swedish Words(1000 Words): Swedish, #0
From Everand
Key & Common Swedish Words A Vocabulary List of High Frequency Swedish Words(1000 Words): Swedish, #0
MostUsedWords
2/5 (4)
Aman Pandey Resume 20241012
No ratings yet
Aman Pandey Resume 20241012
2 pages
BMW 5 Series BSI BRI Sheet
No ratings yet
BMW 5 Series BSI BRI Sheet
1 page
Unit 2
No ratings yet
Unit 2
18 pages
Tekstong Deskriptibo
No ratings yet
Tekstong Deskriptibo
1 page
TRAINEE's PROGRESS SHEET-TDNC2-JB - RAMOS
No ratings yet
TRAINEE's PROGRESS SHEET-TDNC2-JB - RAMOS
3 pages
Exp22 Excel Ch04 CumulativeAssessment Variation Rockville Auto Sales Instructions
No ratings yet
Exp22 Excel Ch04 CumulativeAssessment Variation Rockville Auto Sales Instructions
2 pages
Ford Truck f650 f750 Wiring Diagrams 1999
No ratings yet
Ford Truck f650 f750 Wiring Diagrams 1999
16 pages
LXV50 2stroke Workshop Manual PDF
No ratings yet
LXV50 2stroke Workshop Manual PDF
162 pages
Lecture 2 - Problem Solving Process
No ratings yet
Lecture 2 - Problem Solving Process
32 pages
Skillnet Ireland - Network Brand Guidelines
100% (1)
Skillnet Ireland - Network Brand Guidelines
59 pages
CG Report Final-Full
No ratings yet
CG Report Final-Full
24 pages
Pathfinder Solution Overview
No ratings yet
Pathfinder Solution Overview
2 pages
Datasheet 1 RTG 1223160 E 2,400.0
No ratings yet
Datasheet 1 RTG 1223160 E 2,400.0
2 pages
Clinical Job Aid Radiant Warmer Phoenix
No ratings yet
Clinical Job Aid Radiant Warmer Phoenix
2 pages
Module 1 - Introduction To Computer Networks
No ratings yet
Module 1 - Introduction To Computer Networks
9 pages
Draft - R1-2312083 Summary of UE Features For NR NTN - v002 - DCM - HW&HiSi
No ratings yet
Draft - R1-2312083 Summary of UE Features For NR NTN - v002 - DCM - HW&HiSi
23 pages
BCN Campus Recruitment Process - FAQ
No ratings yet
BCN Campus Recruitment Process - FAQ
1 page
Physics Investigatory Project
No ratings yet
Physics Investigatory Project
17 pages
Rigging Safety
100% (1)
Rigging Safety
27 pages
GetTempFileName Function (Winbase.h) - Win32 Apps - Microsoft Learn
No ratings yet
GetTempFileName Function (Winbase.h) - Win32 Apps - Microsoft Learn
4 pages
1-ICT Topic 3
100% (1)
1-ICT Topic 3
6 pages
IT 2023 - Digital - (SEGi Susan 012-2820 251)
No ratings yet
IT 2023 - Digital - (SEGi Susan 012-2820 251)
24 pages
Standard Truss Garage Plan
No ratings yet
Standard Truss Garage Plan
12 pages
System Requirements Guidelines NX 8 5
No ratings yet
System Requirements Guidelines NX 8 5
3 pages
Bricks
No ratings yet
Bricks
34 pages
Fake Snapchat Chat Generator
No ratings yet
Fake Snapchat Chat Generator
1 page
NIJ-0108.01 Ballistic Resistant Protective Materials
100% (1)
NIJ-0108.01 Ballistic Resistant Protective Materials
16 pages
Hardness Shore A vs. Shore D - Darwin Microfluidics
No ratings yet
Hardness Shore A vs. Shore D - Darwin Microfluidics
2 pages

Topic 4 W4 - Text Processing

Uploaded by

Topic 4 W4 - Text Processing

Uploaded by

Search Engines

Information Retrieval in Practice

All slides ©Addison Wesley, 2008

• How many pages contain all of the query terms?

• Assuming that terms occur independently

Collection size (N) is 25,205,179

Words with the Arabic root ktb

You might also like