Analysis of large bodies of textual data
Abstract
In various example embodiments, a textual identification system is configured to receive a set of search terms and identify a set of textual data based on the search terms. The textual identification system retrieves a data structure including textual identifications for the set of textual data and processes the data structure to generate a modified data structure. The textual identification system sums rows within the modified data structure and identifies one or more elements of interest. The textual identification system then causes presentation of the elements of interest in a first portion of a graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A method, comprising:
accessing, by one or more processors of a machine, a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets; dynamically partitioning, by one or more processors, the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with one or more search terms; identifying, by one or more processors, a set of textual data based on the one or more search terms; retrieving, by one or more processors, a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data; processing, by the one or more processors, the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms; summing rows, by the one or more processors, within the modified data structure, the rows including values for data elements included in each of the identified set of textual data; identifying, by the one or more processors, one or more elements of interest within the set of textual data based on the summed rows of the modified data structure; and causing, by one or more processors, presentation of the elements of interest in a first portion of a graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.
3 . The method of claim 2 , wherein the set of textual corpora comprises a set of document corpora, each textual corpus from the set of document corpora comprises a document corpus, and the set of text sets comprises a set of documents.
4 . The method of claim 2 , wherein each textual corpus is associated with a separate data source.
5 . The method of claim 2 , wherein each textual corpus is associated with a distinct client device.
6 . The method of claim 2 , wherein the causing the presentation of the elements of interest and the textual identifications comprises:
causing presentation of at least a portion of a text set of the set of textual data in a third portion of the graphical user interface.
7 . The method of claim 2 , wherein the causing the presentation of the elements of interest and the textual identifications comprises:
identifying an element type for each element in the elements of interest; and causing presentation of a visual indicator differentiating the elements of interest based on an element type.
8 . The method of claim 2 , further comprising:
determining a context of occurrence for each element of interest within the set of textual data; in response to determining the context of occurrence for each element of interest, generating a set of tokens for each element of interest, the set of tokens representing the context of occurrence; identifying an overlap of two or more of the elements of interest based on the set of tokens for the two or more elements of interest; and linking two or more elements of interest.
9 . The method of claim 2 , further comprising:
receiving a selection of a given textual corpus from the set of textual corpora, the set of textual data identified from the selected given textual corpus.
10 . A computer implemented system, comprising:
one or more processors; and a processor-readable storage device comprising processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
accessing a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets;
dynamically partitioning the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with one or more search terms;
identifying a set of textual data based on the one or more search terms;
retrieving a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data;
processing the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms;
summing rows within the modified data structure, the rows including values for data elements included in each of the identified set of textual data;
identifying one or more elements of interest within the set of textual data based on the summed rows of the modified data structure; and
causing presentation of the elements of interest in a first portion of a graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.
11 . The system of claim 10 , wherein the set of textual corpora comprises a set of document corpora, each textual corpus from the set of document corpora comprises a document corpus, and the set of text sets comprises a set of documents.
12 . The system of claim 10 , wherein each textual corpus is associated with a separate data source.
13 . The system of claim 10 , wherein each textual corpus is associated with a distinct client device.
14 . The system of claim 10 , wherein the causing the presentation of the elements of interest and the textual identifications comprises:
causing presentation of at least a portion of a text set of the set of textual data in a third portion of the graphical user interface.
15 . The system of claim 10 , wherein the causing the presentation of the elements of interest and the textual identifications comprises:
identifying an element type for each element in the elements of interest; and causing presentation of a visual indicator differentiating the elements of interest based on an element type.
16 . The system of claim 10 , wherein the operations further comprise:
determining a context of occurrence for each element of interest within the set of textual data; in response to determining the context of occurrence for each element of interest, generating a set of tokens for each element of interest, the set of tokens representing the context of occurrence; identifying an overlap of two or more of the elements of interest based on the set of tokens for the two or more elements of interest; and linking two or more elements of interest.
17 . The system of claim 10 , wherein the operations further comprise:
receiving a selection of a given textual corpus from the set of textual corpora, the set of textual data identified from the selected given textual corpus.
18 . A processor-readable storage device comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
accessing a set of textual corpora, each textual corpus of the set of textual corpora containing one or more text sets; dynamically partitioning the set of textual corpora to identify a textual corpus from the set of textual corpora containing the set of text sets associated with one or more search terms; identifying a set of textual data based on the one or more search terms; retrieving a data structure including textual identifications for the set of textual data and an indication of one or more data elements within one or more text sets of the set of textual data; processing the data structure to generate a modified data structure, the modified data structure generated by reducing to text sets included in the set of textual data identified based on the one or more search terms; summing rows within the modified data structure, the rows including values for data elements included in each of the identified set of textual data; identifying one or more elements of interest within the set of textual data based on the summed rows of the modified data structure; and causing presentation of the elements of interest in a first portion of a graphical user interface and the textual identifications for the set of textual data in a second portion of the graphical user interface.
19 . The processor-readable storage device of claim 18 , wherein the set of textual corpora comprises a set of document corpora, each textual corpus from the set of document corpora comprises a document corpus, and the set of text sets comprises a set of documents.
20 . The processor-readable storage device of claim 18 , wherein each textual corpus is associated with a separate data source.
21 . The processor-readable storage device of claim 18 , wherein each textual corpus is associated with a distinct client device.Join the waitlist — get patent alerts
Track US2019243897A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.