Identifying key phrases within documents
Abstract
Systems are used for identifying key phrases within documents. These systems utilize a tags and a tag index to determine what a document primarily relates to. For example, an integrated data flow and extract-transform-load pipeline, crawls, parses and word breaks large corpuses of documents in database tables. Documents can be broken into tuples. The tuples can be sent to a heuristically based algorithm that uses statistical language models and weight plus cross-entropy threshold functions to summarize the document into its “top N” most statistically significant phrases. These systems can scale efficiently (e.g., linearly) and (potentially large numbers of) documents can be characterized by salient and relevant key phrases (tags).
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . At a computing system including one or more processors and system memory, a method implemented by the computing system for identifying key phrases within a document, the method comprising:
an act of accessing a document containing a plurality of textual phrases; for each textual phrase in the plurality of textual phrases contained the document:
an act of generating a location list for the textual phrase, the location list indicating one or more locations of the textual phrase within the document;
an act of assigning a score to the textual phrase based on the contents of the location list for the textual phrase relative to the occurrence of the textural phrase in a training set of data;
an act of ranking the plurality of textual phrases according to the assigned scores; an act of selecting a subset of the plurality of textual phrases from within the document based on rankings; and an act of populating a key phrase data structure the selected subset of the plurality of textual phrases.
2 . The method as recited in claim 1 , further comprising an act of generating the training set of data through a plurality of queries to a search engine.
3 . The method as recited in claim 1 , wherein the training set of data is a language model.
4 . The method as recited in claim 1 , wherein the ranking includes sorting a plurality of textual phrases associated with the document based on assigned scores such that textual phrases that are determined to have a similar relevancy to the document are grouped together.
5 . The method as recited in claim 1 , wherein the method further includes appending the document with a document identifier.
6 . The method as recited in claim 5 , wherein the document identifier indicates where the textual phrase occurs in the document.
7 . The method as recited in claim 1 , wherein the method includes identifying a set of one or more most statistically significant textual phrases in the document.
8 . A computer program product for use at a computing system, the computer program product comprising one or more computer storage devices having stored thereon computer-executable instructions that, when executed at a processor, cause the computing system to perform a method for identifying key phrases within a document, wherein the method includes the computing system performing the following:
an act of accessing a document containing a plurality of textual phrases; for each textual phrase in the plurality of textual phrases contained the document:
an act of generating a location list for the textual phrase, the location list indicating one or more locations of the textual phrase within the document;
an act of assigning a score to the textual phrase based on the contents of the location list for the textual phrase relative to the occurrence of the textural phrase in a training set of data;
an act of ranking the plurality of textual phrases according to the assigned scores; an act of selecting a subset of the plurality of textual phrases from within the document based on rankings; and an act of populating a key phrase data structure the selected subset of the plurality of textual phrases.
9 . The computer program product as recited in claim 8 , further comprising an act of generating the training set of data through a plurality of queries to a search engine.
10 . The computer program product as recited in claim 8 , wherein the training set of data is a language model.
11 . The computer program product as recited in claim 8 , wherein the ranking includes sorting a plurality of textual phrases associated with the document based on assigned scores such that textual phrases that are determined to have a similar relevancy to the document are grouped together.
12 . The computer program product as recited in claim 8 , wherein the method further includes appending the document with a document identifier and wherein the document identifier indicates where the textual phrase occurs in the document.
13 . The computer program product as recited in claim 8 , wherein the method includes identifying a set of one or more statistically significant textual phrases in the document and using a tag index to identify what the document primarily relates to based on the one or more most statistically significant textual phrases in the document.
14 . A computing system comprising:
at least one processor; and one or more computer-readable media having stored computer-executable instructions that, when executed by the at least one processor, cause the computing system to perform a method for identifying key phrases within a document, wherein the method includes the computing system performing the following:
an act of accessing a document containing a plurality of textual phrases;
for each textual phrase in the plurality of textual phrases contained the document:
an act of generating a location list for the textual phrase, the location list indicating one or more locations of the textual phrase within the document;
an act of assigning a score to the textual phrase based on the contents of the location list for the textual phrase relative to the occurrence of the textural phrase in a training set of data;
an act of ranking the plurality of textual phrases according to the assigned scores;
an act of selecting a subset of the plurality of textual phrases from within the document based on rankings; and
an act of populating a key phrase data structure the selected subset of the plurality of textual phrases.
15 . The computing system as recited in claim 14 , further comprising an act of generating the training set of data through a plurality of queries to a search engine.
16 . The computing system as recited in claim 14 , wherein the training set of data is a language model.
17 . The computing system as recited in claim 14 , wherein the ranking includes sorting a plurality of textual phrases associated with the document based on assigned scores such that textual phrases that are determined to have a similar relevancy to the document are grouped together.
18 . The computing system as recited in claim 14 , wherein the method further includes appending the document with a document identifier and wherein the document identifier indicates where the textual phrase occurs in the document.
19 . The computing system as recited in claim 14 , wherein the method includes identifying a set of one or more statistically significant textual phrases in the document and using a tag index to identify what the document primarily relates to based on the one or more most statistically significant textual phrases in the document.
20 . The computing system as recited in claim 14 , wherein the one or more computer-readable media comprises system memory.Join the waitlist — get patent alerts
Track US2013246386A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.