Method for extracting a compact representation of the topical content of an electronic text
Abstract
An electronic document is parsed to remove irrelevant text and to identify the significant elements of the retained text. The elements are assigned scores representing their significance to the topical content of the document. A matrix of element-pairs is constructed such that the matrix nodes represent the result of one or more functions of the scores and other attributes of the paired elements. The resulting matrix is a compact representation of topical content that affords great precision in information retrieval applications that depend on measurements of the relatedness of topical content.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method for extracting topical content from an electronic document, comprising:
identifying elements from the document, wherein elements are words or phrases that appear in the document; calculating a score for each element, wherein the score is related to the element's significance in the document; calculating a sentential affinity value for each element pair having non-redundant elements, wherein the sentential affinity value is at least dependent upon the scores of the elements of the element pair and a sentence radius, wherein the sentence radius is a number of sentences separating closest occurrences of the elements of the element pair in the document; storing the sentential affinity values and associated element pair in a memory as a data structure, wherein the data structure is executable by information retrieval systems.
2 . The method of claim 1 , further comprising eliminating sections of the document determined to contribute insignificantly to a topical content of the document.
3 . The method of claim 1 wherein the document includes more than one section, and further wherein identifying elements from the document comprises:
identifying elements from each section of the document; eliminating redundant elements to create a global list of elements.
4 . The method of claim 3 wherein calculating a score for each element includes combining section scores of a redundant element.
5 . The method of claim 1 , further comprising determining a sentence set for each element, wherein each sentence in the document is sequentially numbered, and the sentence set comprises ordinal numbers associated with each sentence in the document in which the element appears.
6 . The method of claim 5 , further comprising biasing the score for each element based upon a topical selection from which the element was extracted, a composition of a sentence set determined for that element, or a comparison of elements to a predetermined list of elements.
7 . The method of claim 1 wherein calculating a sentential affinity value comprises calculating an average of the scores of the elements of the element pair.
8 . The method of claim 7 wherein calculating a sentential affinity value further comprises weighting the average of the scores of the elements of the element pair based upon the sentence radius.
9 . The method of claim 1 wherein the data structure is a matrix or a vector.
10 . The method of claim 1 , further comprising ordering the stored element pairs based upon the calculated sentential affinity values.
11 . The method of claim 1 , further comprising truncating the data structure to a configurable number of dimensions.
12 . A computer-implemented method of extracting topical content from an electronic document, comprising:
identifying elements in the document, wherein elements are words or tokens that appear in the document; calculating a score for each element, wherein the score is related to the element's significance to the document; calculating a sentential affinity value for each element pair having non-redundant elements, wherein the sentential affinity value is at least dependent upon the scores of the elements of the element pair and one or more attributes of the element pair; storing the sentential affinity values and associated element pairs in a memory as a data structure, wherein the data structure is executable by information retrieval systems.
13 . The method of claim 12 wherein an attribute of the element pair includes an index of earliest occurrence for each of the elements of the element pair.
14 . The method of claim 12 wherein an attribute of the element pair includes a sentence radius, wherein the sentence radius is a number of sentences separating closest occurrences of the elements of the element pair in the document.
15 . The method of claim 12 wherein an attribute of the element pair includes a sentence set for each element, wherein each sentence in the document is sequentially numbered, and the sentence set comprises ordinal numbers associated with each sentence in the document in which the element appears.
16 . The method of claim 15 , further comprising biasing the score for each element based upon a topical selection from which the element was extracted, a composition of a sentence set for that element, or a comparison of elements to a predetermined list of elements.
17 . The method of claim 12 wherein identifying elements from the document includes parsing the document to remove irrelevant text and eliminating text not representative of topical content of the document.
18 . The method of claim 12 wherein calculating a score for each element comprises using a normalized, relative frequency measure.
19 . The method of claim 12 , further comprising truncating the data structure to a configurable number of dimensions.
20 . The method of claim 12 , further comprising ordering the stored element pairs based upon the calculated sentential affinity values.
21 . A computer-implemented method for extracting topical content from an electronic document, comprising:
identifying elements from the document, wherein elements are words or phrases that appear in the document, further wherein the document includes more than one section, and identifying elements from the document comprises:
identifying elements from each section of the document; eliminating redundant elements to create a global list of elements;
eliminating sections of the document determined to contribute insignificantly to a topical content of the document;
calculating a score for each element, wherein the score is related to the element's significance in the document, and calculating a score for each element includes combining section scores of a redundant element;
biasing the score for each element based upon a topical selection from which the element was extracted, a composition of a sentence set determined for that element, or a comparison of elements to a predetermined list of elements;
determining a sentence set for each element, wherein each sentence in the document is sequentially numbered, and the sentence set comprises ordinal numbers associated with each sentence in the document in which the element appears;
calculating a sentential affinity value for each element pair having non-redundant elements, wherein the sentential affinity value is at least dependent upon the scores of the elements of the element pair and a sentence radius, and the sentence radius is a number of sentences separating closest occurrences of the elements of the element pair in the document, and calculating a sentential affinity value comprises calculating an average of the scores of the elements of the element pair and weighting the average of the scores of the elements of the element pair based upon the sentence radius;
storing the sentential affinity values and associated element pair in a memory as a data structure, wherein the data structure is executable by information retrieval systems, and the data structure is a matrix or a vector;
ordering the stored element pairs based upon the calculated sentential affinity values;
truncating the data structure to a configurable number of dimensions.
22 . A computer-implemented method of extracting topical content from an electronic document, comprising:
identifying elements in the document, wherein elements are words or tokens that appear in the document, and identifying elements from the document includes parsing the document to remove irrelevant text and eliminating text not representative of topical content of the document; calculating a score for each element, wherein the score is related to the element's significance to the document, and calculating a score for each element comprises using a normalized, relative frequency measure; biasing the score for each element based upon a topical selection from which the element was extracted, a composition of a sentence set for that element, or a comparison of elements to a predetermined list of elements; calculating a sentential affinity value for each element pair having non-redundant elements, wherein the sentential affinity value is at least dependent upon the scores of the elements of the element pair and one or more attributes of the element pair, and further wherein an attribute of the element pair includes an index of earliest occurrence for each of the elements of the element pair; a sentence radius, wherein the sentence radius is a number of sentences separating closest occurrences of the elements of the element pair in the document; and a sentence set for each element, wherein each sentence in the document is sequentially numbered, and the sentence set comprises ordinal numbers associated with each sentence in the document in which the element appears; storing the sentential affinity values and associated element pairs in a memory as a data structure, wherein the data structure is executable by information retrieval systems; truncating the data structure to a configurable number of dimensions; ordering the stored element pairs based upon the calculated sentential affinity values.Join the waitlist — get patent alerts
Track US2009292698A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.