US2006212421A1PendingUtilityA1
Contextual phrase analyzer
Individually held — no corporate assignee on recordPriority: Mar 18, 2005Filed: Mar 13, 2006Published: Sep 21, 2006
Est. expiryMar 18, 2025(expired)· nominal 20-yr term from priority
Inventors:Guillermo Oyarce
G06F 16/36G06F 16/958
15
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method and a computer system for implementing a contextual phrase analyzer engine are provided. The method includes selecting at least one of a plurality of document frequencies associated with a plurality of words used in a plurality of documents and selecting a subset of the plurality of words based on the at least one selected document frequency. The method also includes selecting at least one of the words in the subset of the plurality of words based on word frequencies associated with each word in the subset of the plurality of words.
Claims
exact text as granted — not AI-modified1 . A method, comprising:
selecting at least one of a plurality of document frequencies associated with a plurality of words used in a plurality of documents; selecting a subset of the plurality of words based on said at least one selected document frequency; and selecting at least one of the words in the subset of the plurality of words based on word frequencies associated with each word in the subset of the plurality of words.
2 . The method of claim 1 , further comprising determining the plurality of document frequencies using information indicative of the plurality of words used in the plurality of documents.
3 . The method of claim 2 , wherein determining the plurality of document frequencies comprises accessing the information indicative of the plurality of words used in the plurality of documents.
4 . The method of claim 1 , wherein selecting at least one of the plurality of document frequencies comprises selecting at least one of the plurality of document frequencies based upon a distribution of document frequencies associated with the plurality of words used in the plurality of documents.
5 . The method of claim 4 , wherein selecting at least one of the plurality of document frequencies comprises rejecting document frequencies at a low document frequency tail of the distribution and a high document frequency tail of the distribution.
6 . The method of claim 4 , wherein selecting at least one of the plurality of document frequencies comprises rejecting document frequencies at the low document frequency tail of the distribution based on a first predetermined parameter and the high document frequency tail of the distribution based on a second predetermined parameter.
7 . The method of claim 1 , wherein selecting the subset of the plurality of words comprises selecting at least one word that appears in the plurality of documents at said at least one document frequency.
8 . The method of claim 1 , wherein selecting at least one of the subset of the plurality of words comprises selecting at least one of the subset of the plurality of words having relatively high word frequencies.
9 . The method of claim 1 , wherein selecting at least one of the subset of the plurality of words comprises selecting at least one of the subset of the plurality of words having a word frequency above a first predetermined word frequency.
10 . The method of claim 9 , wherein selecting at least one of the subset of the plurality of words comprises selecting at least one of the subset of the plurality of words having a word frequency below a second predetermined word frequency.
11 . The method of claim 1 , further comprising providing information indicative of said at least one word selected from the subset of the plurality of words to a user via a user interface.
12 . A computer system, comprising:
at least one processing unit configured to:
select at least one of a plurality of document frequencies associated with a plurality of words used in a plurality of documents;
select a subset of the plurality of words based on said at least one selected document frequency; and
select at least one of the words in the subset of the plurality of words based on word frequencies associated with each word in the subset of the plurality of words.
13 . The computer system of claim 12 , wherein the processing unit is configured to determine the plurality of document frequencies using information indicative of the plurality of words used in the plurality of documents.
14 . The computer system of claim 13 , further comprising at least one memory unit, and wherein the processing unit is configured to access information indicative of the plurality of words used in the plurality of documents from the memory unit.
15 . The computer system of claim 12 , wherein the processing unit is configured to select at least one of the plurality of document frequencies based upon a distribution of document frequencies associated with the plurality of words used in the plurality of documents.
16 . The computer system of claim 15 , wherein the processing unit is configured to reject document frequencies at a low document frequency tail of the distribution and a high document frequency tail of the distribution.
17 . The computer system of claim 16 , wherein the processing unit is configured to reject document frequencies at the low document frequency tail of the distribution based on a first predetermined parameter and the high document frequency tail of the distribution based on a second predetermined parameter.
18 . The computer system of claim 17 , wherein the processing unit is configured to select at least one word that appears in the plurality of documents at said at least one document frequency.
19 . The computer system of claim 12 , wherein the processing unit is configured to select at least one of the subset of the plurality of words having relatively high word frequencies.
20 . The computer system of claim 12 , wherein the processing unit is configured to select at least one of the subset of the plurality of words having a word frequency above a first predetermined word frequency.
21 . The computer system of claim 20 , wherein the processing unit is configured to select at least one of the subset of the plurality of words having a word frequency below a second predetermined word frequency.
22 . The computer system of claim 12 , further comprising a display unit configured to display information indicative of said at least one word selected from the subset of the plurality of words to a user via a user interface.Join the waitlist — get patent alerts
Track US2006212421A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.