US2022198152A1PendingUtilityA1
Generating hypotheses in data sets
Est. expiryJan 15, 2034(~7.5 yrs left)· nominal 20-yr term from priority
G06N 5/01G06N 7/01G06N 3/126G06N 5/022G06F 40/40G06N 7/005G06N 5/003
64
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method for generating hypotheses in a corpus of data comprises selecting a form of ontology; coding the corpus of data based on the form of the ontology; generating ontology space based on coding results and the ontology; transforming the ontology space into a hypothesis space by grouping hypotheses; weighing hypotheses included in the hypothesis space; and applying a science-based optimization algorithm configured to model a science-based treatment of the weighted hypotheses.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A method of identifying hypotheses in a corpus of data, the method comprising:
receiving an ontology by one or more computers, the ontology including a plurality of fields and a plurality of choices for each of the fields such that the ontology includes a plurality of ontology vectors that each include one choice for each of the fields of the ontology, the ontology vectors being organized as a multi-dimensional space wherein:
each dimension of the multi-dimensional space represents one or more fields of the ontology; and
ontology vectors representing similar and/or related concepts are closer together in the multi-dimensional space than ontology vectors representing and/or unrelated concepts;
receiving the corpus of data by the one or more computers; identifying ontology vectors in the corpus of data, by the one or more computers, by detecting data in the corpus of data that corresponds to ontology vectors in the ontology; grouping the identified ontology vectors that describe similar and/or related concepts into groups, by the one more computers, each group of similar and/or related ontology vectors representing a hypothesis; weighing each of the hypotheses by the one or more computers; and applying an optimization algorithm, by the one or more computers, to rank the hypotheses in accordance with the weight of each hypothesis.
22 . The method of claim 21 , wherein the optimization algorithm comprises one of a simulated annealing algorithm, a Monte Carlo-based algorithm, or a genetic algorithm.
23 . The method of claim 21 , wherein the optimization algorithm ranks the hypotheses in accordance with the weight of each hypothesis by:
ranking the deepest troughs of a multi-dimensional surface having troughs that each represent a group of similar and/or related ontology vectors and each have a depth proportional to the weight of the group of ontology vectors; or ranking the fittest population members of a population having population members that each represent a group of similar and/or related ontology vectors and each have a fitness proportional to the weight of the group of ontology vectors.
24 . The method of claim 21 , wherein the weight each of hypothesis is based on frequency of one or more words, parts of speech, thresholding of concepts, or exclusions.
25 . The method of claim 21 , wherein:
the corpus of data includes a plurality of documents; the method further comprises weighting one or more of the documents; and the weight each of hypothesis is based at least in part on the weight of one or more documents having data corresponding to the ontology vector representing each hypothesis.
26 . The method of claim 24 , wherein the weight of each document is based on a source of capture of the document, volume of the document, uniqueness of the document, or variance of the document.
27 . The method of claim 21 , wherein:
the form of the ontology includes a weight for each of the fields; and the weight of each hypothesis is based at least in part on the weight of the fields of the ontology vector representing each hypothesis.
28 . The method of claim 21 , wherein:
the ontology includes N fields; and the multi-dimensional space includes N dimensions, each of the N dimensions representing one of the N fields of the ontology.
29 . The method of claim 21 , wherein:
the ontology includes N fields; grouping the identified ontology vectors comprises separating the N fields of the ontology into R groups; and the multi-dimensional space includes R dimensions, each of the R dimensions representing one of the R groups.
30 . The method of claim 21 , wherein the identified ontology vectors that describe similar and/or related concepts are grouped using one or more clustering techniques.
31 . The method of claim 30 , wherein the one or more clustering techniques include hierarchies, filters and thresholds, topic models, or conditional random fields.
32 . The method of claim 21 , wherein the optimization algorithm de-weights trivial or uninteresting hypotheses by:
introducing a random variation or mutation into data representing the groups of similar and/or related ontology vectors; and determining an anticipation level of each group of ontology vectors.
33 . The method of claim 32 , wherein:
the optimization algorithm comprises a simulated annealing algorithm; and the anticipation level of each group of ontology vectors is determined based on a slope of descent or accent of a local minima representing the group of ontology vectors.
34 . The method of claim 32 , wherein:
the optimization algorithm comprises a genetic algorithm; the anticipation level of each group of ontology vectors is determined based on a fitness level of a population member representing the group of ontology vectors.
35 . The method of claim 21 , wherein weighting and ranking the hypotheses comprises:
storing personalized criteria of a user; and filtering the hypotheses to de-weight hypotheses that are trivial or uninteresting to the user.
36 . The method of claim 35 , wherein the personalized criteria is determined based on hypotheses previously considered by the user.
37 . The method of claim 35 , further comprising:
storing the identity or role of the user, wherein the personalized criteria is determined based on the identity or role of the user.
38 . The method of claim 21 , wherein the hypotheses are ranked based on the path through the multi-dimensional space by which the group of ontology vectors representing each hypothesis was discovered by the optimization algorithm.
39 . The method of claim 21 , wherein the hypotheses are ranked in a stateless manner based on the positions of the groups of ontology vectors representing each hypotheses in the multi-dimensional space.
40 . The method of claim 21 , further comprising:
outputting at least some of the ranked hypotheses for display to a user.
41 . A system for identifying hypotheses in a corpus of data, the method comprising:
non-transitory computer readable storage media that stores the corpus of data; a content server that:
receives an ontology by one or more computers, the ontology including a plurality of fields and a plurality of choices for each of the fields such that the ontology includes a plurality of ontology vectors that each include one choice for each of the fields of the ontology, the ontology vectors being organized as a multi-dimensional space wherein:
each dimension of the multi-dimensional space represents one or more fields of the ontology; and
ontology vectors representing similar and/or related concepts are closer together in the multi-dimensional space than ontology vectors representing dissimilar and/or unrelated concepts;
identifies ontology vectors in the corpus of data;
groups the identified ontology vectors that describe similar and/or related concepts into groups, each group of similar and/or related ontology vectors representing a hypothesis;
weighs each of the hypotheses; and
using an optimization algorithm to rank the hypotheses in accordance with the weight of each hypothesis.Join the waitlist — get patent alerts
Track US2022198152A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.