Method and system to provide related data
Abstract
Methods and systems of providing related information to a source document are described. The method may include accessing the source document displayed to a user in a graphical user interface (GUI) of a client device. The source document includes numerical data and text. Discovered data corresponding to the numerical data included in the source document is then identified. Further, a database trained with a machine-learning algorithm to identify time series data related data associated with the text is accessed. The discovered data with a discovered data identifier and the time series related data is then displayed in the GUI. In example embodiments, the methods and systems described herein interact with applications such as spreadsheets applications, email clients, word processing applications, webpages and the like.
Claims
exact text as granted — not AI-modified1 . (canceled)
2 . A method of providing information related to a source document, the method comprising:
receiving the source document from a client device via a communication network; accessing, using one or more hardware processors, the source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of a client device; generating, using the one or more hardware processors, discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document;
accessing, using the one or more hardware processors, a database trained with a machine learning algorithm to identify time series related data associated with the text; and
communicating the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.
3 . The method of claim 2 , wherein the accessing of the source document, the generating the discovered data, and the accessing the database occurs automatically on the fly without user selection.
4 . The method of claim 2 , wherein the GUI comprises:
a document zone displaying the source document; a discovered data display zone to display the numerical data communicated to the client device and the discovered data identifier communicated to the client device; and a related data display zone to display the time series related data communicated to the client device.
5 . The method of claim 2 , further comprising preprocessing the source document using a natural language processing algorithm.
6 . The method of claim 2 , wherein the generating the discovered data further comprises:
accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and generating the discovered data based on at least the machine learning model, the numerical data, and the text.
7 . The method of claim 2 , wherein the identifying the time series related data comprises:
accessing data in the machine learning model; accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and generating the time series related data based on both the machine learning model and the numerical data and text from the source document.
8 . The method of claim 2 , wherein the time series related data is displayed in one or more graphs in the GUI of the client device.
9 . The method of claim 2 , wherein the generating the discovered data further comprises:
searching for similar relations for named entities based on the named entities derived from the source document, a syntax tree and a dependency tree derived from the source document, and a relation extraction model; classifying at least some of the similar relations; and converting the classified relations to define the discovered data.
10 . The method of claim 2 , wherein the accessing the database trained with the machine learning algorithm to identify time series related data associated with the text further comprises:
identifying primary words from sentences extracted from the source document; indexing terms of the text; identifying terms from the indexed terms and the primary words to obtain a term set; transitioning the terms set to a series set; and generating related data based on relevance of the series set.
11 . The method of claim 2 , further comprising parsing the source document for key values corresponding to reference values provided in a data repository.
12 . The method of claim 2 , wherein the database is remotely located from the client device, the method further comprising accessing the database via a network to identify the time series related data associated with the text;
receiving the discovered data with the discovered data identifier and the time series related data via the network; and displaying the discovered data with the discovered data identifier and the time series related data in the GUI.
13 . The method of claim 2 , wherein the GUI of the client device is presented in a web browser, the method further comprising:
providing a plurality of hyperlinks in a webpage associated with the discovered data and the time series related data; monitoring selection of a hyperlink of the plurality of hyperlinks; and communicating further related data to the client device, via the communication network, upon selection of the hyperlink.
14 . The method of claim 2 , wherein the method is at least partially performed by a plug-in specially configured to interact with an application displaying the source document.
15 . The method of claim 2 , wherein the source document is displayed in an application selected from a group consisting of a web browser, a spreadsheet application, a word processing application, and an email client.
16 . A computerized system comprising:
a receiving module implemented by one or more hardware processors and configured to receive a source document from a client device via a communication network an access module implemented by the one or more hardware processors and configured to access a source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of the client device; a discovered data module implemented by the one or more hardware processors and configured to generate discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document; a database access module implemented by the one or more hardware processors and configured to access a database trained with a machine learning algorithm to identify time series related data associated with the text; and a display module configured to communicate the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.
17 . The computerized system of claim 16 , wherein the GUI comprises:
a document zone displaying the source document; a discovered data display zone to display the numerical data communicated to the client device and the discovered data identifier communicated to the client device; and a related data display zone to display the time series related data communicated to the client device.
18 . The computerized system of claim 16 , wherein the generating the discovered data further comprises:
accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and generating the discovered data based on at least the machine learning model, the numerical data, and the text.
19 . The computerized system of claim 16 , wherein the identifying the time series related data comprises:
accessing data in the machine learning model; accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and generating the time series related data based on both the machine learning model and the numerical data and text from the source document.
20 . The computerized system of claim 16 , wherein the generating the discovered data further comprises:
searching for similar relations for named entities based on the named entities derived from the source document, a syntax tree and a dependency tree derived from the source document, and a relation extraction model; classifying at least some of the similar relations; and converting the classified relations to define the discovered data.
21 . A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
receiving a source document from a client device via a communication network; accessing, using one or more hardware processors, the source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of a client device; generating, using the one or more hardware processors, discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document; accessing, using the one or more hardware processors, a database trained with a machine learning algorithm to identify time series related data associated with the text; and communicating the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.Join the waitlist — get patent alerts
Track US2019034835A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.