US2009083027A1PendingUtilityA1
Automatic text skimming using lexical chains
Individually held — no corporate assignee on recordPriority: Aug 16, 2007Filed: Aug 15, 2008Published: Mar 26, 2009
Est. expiryAug 16, 2027(~1 yrs left)· nominal 20-yr term from priority
Inventors:William A. Hollingsworth
G06F 40/30G06F 16/31G06F 40/284
48
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Automatic text skimming using lexical chains may be provided. First, at least one lexical chain may be created from an electronic document. Next, a list of positions within the electronic document may be created. The positions may include where at least one concept represented by one of the at least one lexical chain is mentioned. In addition, a list of the position where the at least one concept is mentioned may be assembled. A selection of at least one concept may be received from the list.
Claims
exact text as granted — not AI-modified1 . A method for automatic text skimming using lexical chains, the method comprising:
creating at least one lexical chain from an electronic document; creating a concept list of positions within the electronic document where at least one concept represented by one of the at least one lexical chain is mentioned; assembling a positions list of positions where the at least one concept is mentioned; receiving a selection of at least one concept from the positions list; and highlighting, in response to receiving the selection, sections of the electronic document that contain the selected at least one concept.
2 . The method of claim 1 , wherein creating the at least one lexical chain from the electronic document comprises determining at least one major topic from the electronic document.
3 . The method of claim 1 , further comprising:
creating an annotated table of contents, wherein creating the annotated table of contents comprises merging an index and a table of contents of the electronic document; and wherein receiving the selection of the at least one concept from the position list comprises selecting the at least one concept from the annotated table of contents.
4 . The method of claim 3 , wherein creating the annotated table of contents, wherein the annotated table of contents is a representation of contents of the electronic document and creating the annotated table of contents comprises:
constructing a list of at least one of the following: chapter titles, section titles, and subsection titles in the electronic document; for each of the chapter titles, the section titles, and the subsection titles, constructing a list of at least one of the following: concepts and lexical chains referred to in each of the chapter titles, the section titles, and the subsection titles; associating the chapter titles, the section titles, and the subsection titles with a portion of the text in the electronic document.
5 . The method of claim 1 , wherein at least one lexical chain is created from a text such that each lexical chain represents a concept discussed in the text.
6 . The method of claim 1 , wherein for each lexical chain creating a list of positions in the text, denoted by at least one of the following: sentence numbers and page numbers, corresponding to a use of the lexical chain.
7 . The method of claim 1 , further comprising using a section heading information, provided by an author of the electronic document, to make a record of how many times the at least one lexical chain is mentioned in the section.
8 . The method of claim 1 , further comprising creating a list of the most used lexical chains for at least one of the following: a chapter, a section, and a subsection.
9 . The method of claim 1 , further comprising:
assembling a chapter list of at least one of a chapter, a section, and a subsection in the electronic document; and assembling a usage list of lexical chains used in each of the chapters, sections, and subsections.
10 . The method of claim 1 , further comprising:
hiding all unhighlighted text; and presenting the user with a summary of the paper.
11 . The method of claim 1 , further comprising transmitting the at least one lexical chain to a remote computing device.
12 . The method of claim 1 , further comprising broadcasting the at least one lexical chain via a speech synthesizer.
13 . A method for using lexical chains to construct a search index for at least one text document for use by a search engine.
14 . The method of claim 13 , further comprising creating an annotated table of contents, including the lexical chains, for every document to be indexed by the search engine.
15 . The method of claim 13 , further comprising, for each of the lexical chains, recording a location of all sentences in the document that reference each lexical chain.
16 . The method of claim 13 , further comprising computing a contribution of each lexical chain to a representation of the document, wherein a lexical chain representing a key concept in the document receives a higher contribution score than a lexical chain representing a minor point in the document.
17 . The method of claim 13 , further comprising recording all technical terms used in any lexical chain in the document in a technical term list.
18 . The method of claim 17 , further comprising:
computing a contribution of each lexical chain to a representation of the document, and computing a term contribution for each technical term to each lexical chain for the document.
19 . The method of claim 13 , further comprising creating a document index including a lexical chain list of the lexical chains created for the document, for each of the lexical chains in the lexical chain list, creating at least one of the following:
a location list of all locations in the document, denoted by at least one of the following: a sentence number and a page number, referencing the lexical chain; a page list of all page numbers in the document that correspond to pages containing a sentence referencing the lexical chain; and a contribution score denoting the contribution of each lexical chain to a representation of the document.
20 . The method of claim 13 , further comprising creating a document index, wherein the document index includes a terms list of terms used in lexical chains for each document in the document index.
21 . The method of claim 13 , further comprising creating a document index, the document index comprising a list of all terms used in any lexical chain for any document in the index, for each term in the index, a sorted list of documents, denoted by document identifiers, containing the term in a lexical chain.
22 . The method of claim 13 , further comprising creating a document index, the document index comprising a list of all single words used in any lexical chain for any document in the index, for each single word in the index, a sorted list of documents, denoted by document identifiers, containing the word in a lexical chain.
23 . A method for providing a search engine:
receiving a query; determining if the query contains at least one multiword term from a terms list of terms in the search index; and for each single word in the query, retrieving from the search index a document list of at least one document that contains the lexical chains that include the word.
24 . The method of claim 23 , further comprising computing the relevance of each document in the document list to the query.
25 . The method of claim 24 , wherein computing the relevance of each document comprises utilizing at least one of the following to compute the relevance: a number of words in the query, a number of words in the query that are used in the document by a lexical chain, for each word in the query that is used in the document by a lexical chain, a contribution of the word to the representation of the document, a number of multiword terms in the index that are in the query, and for each multiword term in the query, the contribution of that term to the representation of the document.
26 . A computer-readable medium which stores a set of instructions which when executed performs a method for automatic text skimming using lexical chains, the method executed by the set of instructions comprising:
creating at least one lexical chain from an electronic document; creating a concept list of positions within the electronic document where at least one concept represented by one of the at least one lexical chain is mentioned; assembling a positions list indicating where the at least one concept is mentioned; receiving a selection of at least one concept from the positions list; and in response to the selection, performing at least one of the following:
(i) hiding sections of the electronic document that do not contain the selected at least one concept,
(ii) highlighting sections of the electronic document that contain the selected at least one concept, and
(iii) both (i) and (ii).
27 . The computer readable medium of claim 26 , further comprising:
creating an annotated table of contents, wherein creating the annotated table of contents comprises merging an index and a table of contents of the electronic document; and wherein receiving the selection of the at least one concept from the positions list comprises selecting the at least one concept from the annotated table of contents.
28 . A system for automatic text skimming using lexical chains, the system comprising:
a memory storage; and a processing unit coupled to the memory storage, wherein the processing unit is operative to:
create an annotated table of contents from an electronic document;
create concept list comprising a list of positions within the electronic document where at least one concept represented by one of at least one lexical chain is mentioned; and
receive a selection of at least one concept from the annotated table of contents.
29 . The system of claim 28 , wherein the processing unit being operative to receive the selection of at least one concept from the annotated table of contents comprises the processing unit being operative to hide sections of the electronic document that do not contain the selected at least one concept.
30 . The system of claim 28 , wherein the processing unit being operative to receive the selection of at least one concept from the positions list comprises the processing unit being operative to highlight sections of the electronic document that contain the selected at least one concept.
31 . The system of claim 28 , wherein the processing unit is operative to create an annotated table of contents, wherein the processing unit being operative to create the annotated table of contents comprises the processing unit being operative to merge a collection of lexical chains and a table of contents of the electronic document.
32 . A method for filtering adjectives from a lexical chain, the method comprising:
receiving the lexical chain comprising an adjective; testing the adjective to determine if the adjective is at least one of the following: a characteristic adjective and a non-characteristic adjective; when the adjective is a non-characteristic adjective, removing the adjective from the lexical chain; and when the adjective is a characteristic adjective, leaving the adjective in the lexical chain.
33 . The method of claim 32 , wherein testing the adjective comprises determining if the adjective is a non-characteristic adjective, wherein a non-characteristic adjective can be removed from the collocation without changing the meaning of the collocation.
34 . The method of claim 32 , wherein determining if the adjective is at least one of the following: the characteristic adjective and non-characteristic adjective comprises determining a gradability property.
35 . The method of claim 32 , wherein determining if the adjective is at least one of the following: the characteristic adjective and non-characteristic adjective comprises determining a nominalization property
36 . The method of claim 32 , wherein the lexical chain comprises a collocation.
37 . The method of claim 36 , further comprising determining if the collocation is a technical term.
38 . The method of claim 32 , further comprising transmitting the at least one lexical chain to a remote computing device.
39 . The method of claim 32 , further comprising broadcasting the at least one lexical chain via a speech synthesizer.
40 . A method for filtering adjectives before or during formation of a lexical chain, the method comprising:
receiving an adjective; testing the adjective to determine if the adjective is at least one of the following: a characteristic adjective and a non-characteristic adjective; when the adjective is a non-characteristic adjective, form the lexical chain without using the adjective; when the adjective is a characteristic adjective, form the lexical chain, wherein the lexical chain comprises the adjective.Join the waitlist — get patent alerts
Track US2009083027A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.