US2019034835A1PendingUtilityA1

Method and system to provide related data

Assignee: KNOEMA CORPPriority: Jul 17, 2015Filed: Oct 2, 2018Published: Jan 31, 2019
Est. expiryJul 17, 2035(~9 yrs left)· nominal 20-yr term from priority
G06N 5/01G06F 40/295G06F 40/169G06F 3/0484G06F 40/18G06N 20/10G06N 5/02G06F 16/338G06F 40/247G06F 40/30G06F 17/278G06N 99/005G06F 17/246G06F 17/2795G06F 17/30696G06F 17/2785G06F 17/241G06N 20/00
45
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems of providing related information to a source document are described. The method may include accessing the source document displayed to a user in a graphical user interface (GUI) of a client device. The source document includes numerical data and text. Discovered data corresponding to the numerical data included in the source document is then identified. Further, a database trained with a machine-learning algorithm to identify time series data related data associated with the text is accessed. The discovered data with a discovered data identifier and the time series related data is then displayed in the GUI. In example embodiments, the methods and systems described herein interact with applications such as spreadsheets applications, email clients, word processing applications, webpages and the like.

Claims

exact text as granted — not AI-modified
1 . (canceled) 
     
     
         2 . A method of providing information related to a source document, the method comprising:
 receiving the source document from a client device via a communication network;   accessing, using one or more hardware processors, the source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of a client device;   generating, using the one or more hardware processors, discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document;
 accessing, using the one or more hardware processors, a database trained with a machine learning algorithm to identify time series related data associated with the text; and 
   communicating the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.   
     
     
         3 . The method of  claim 2 , wherein the accessing of the source document, the generating the discovered data, and the accessing the database occurs automatically on the fly without user selection. 
     
     
         4 . The method of  claim 2 , wherein the GUI comprises:
 a document zone displaying the source document;   a discovered data display zone to display the numerical data communicated to the client device and the discovered data identifier communicated to the client device; and   a related data display zone to display the time series related data communicated to the client device.   
     
     
         5 . The method of  claim 2 , further comprising preprocessing the source document using a natural language processing algorithm. 
     
     
         6 . The method of  claim 2 , wherein the generating the discovered data further comprises:
 accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and   generating the discovered data based on at least the machine learning model, the numerical data, and the text.   
     
     
         7 . The method of  claim 2 , wherein the identifying the time series related data comprises:
 accessing data in the machine learning model;   accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and   generating the time series related data based on both the machine learning model and the numerical data and text from the source document.   
     
     
         8 . The method of  claim 2 , wherein the time series related data is displayed in one or more graphs in the GUI of the client device. 
     
     
         9 . The method of  claim 2 , wherein the generating the discovered data further comprises:
 searching for similar relations for named entities based on the named entities derived from the source document, a syntax tree and a dependency tree derived from the source document, and a relation extraction model;   classifying at least some of the similar relations; and   converting the classified relations to define the discovered data.   
     
     
         10 . The method of  claim 2 , wherein the accessing the database trained with the machine learning algorithm to identify time series related data associated with the text further comprises:
 identifying primary words from sentences extracted from the source document;   indexing terms of the text;   identifying terms from the indexed terms and the primary words to obtain a term set;   transitioning the terms set to a series set; and   generating related data based on relevance of the series set.   
     
     
         11 . The method of  claim 2 , further comprising parsing the source document for key values corresponding to reference values provided in a data repository. 
     
     
         12 . The method of  claim 2 , wherein the database is remotely located from the client device, the method further comprising accessing the database via a network to identify the time series related data associated with the text;
 receiving the discovered data with the discovered data identifier and the time series related data via the network; and   displaying the discovered data with the discovered data identifier and the time series related data in the GUI.   
     
     
         13 . The method of  claim 2 , wherein the GUI of the client device is presented in a web browser, the method further comprising:
 providing a plurality of hyperlinks in a webpage associated with the discovered data and the time series related data;   monitoring selection of a hyperlink of the plurality of hyperlinks; and   communicating further related data to the client device, via the communication network, upon selection of the hyperlink.   
     
     
         14 . The method of  claim 2 , wherein the method is at least partially performed by a plug-in specially configured to interact with an application displaying the source document. 
     
     
         15 . The method of  claim 2 , wherein the source document is displayed in an application selected from a group consisting of a web browser, a spreadsheet application, a word processing application, and an email client. 
     
     
         16 . A computerized system comprising:
 a receiving module implemented by one or more hardware processors and configured to receive a source document from a client device via a communication network   an access module implemented by the one or more hardware processors and configured to access a source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of the client device;   a discovered data module implemented by the one or more hardware processors and configured to generate discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document;   a database access module implemented by the one or more hardware processors and configured to access a database trained with a machine learning algorithm to identify time series related data associated with the text; and   a display module configured to communicate the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.   
     
     
         17 . The computerized system of  claim 16 , wherein the GUI comprises:
 a document zone displaying the source document;   a discovered data display zone to display the numerical data communicated to the client device and the discovered data identifier communicated to the client device; and   a related data display zone to display the time series related data communicated to the client device.   
     
     
         18 . The computerized system of  claim 16 , wherein the generating the discovered data further comprises:
 accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and   generating the discovered data based on at least the machine learning model, the numerical data, and the text.   
     
     
         19 . The computerized system of  claim 16 , wherein the identifying the time series related data comprises:
 accessing data in the machine learning model;   accessing sentences including extracted from the source document, the sentences including the numerical data and the text; and   generating the time series related data based on both the machine learning model and the numerical data and text from the source document.   
     
     
         20 . The computerized system of  claim 16 , wherein the generating the discovered data further comprises:
 searching for similar relations for named entities based on the named entities derived from the source document, a syntax tree and a dependency tree derived from the source document, and a relation extraction model;   classifying at least some of the similar relations; and   converting the classified relations to define the discovered data.   
     
     
         21 . A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
 receiving a source document from a client device via a communication network;   accessing, using one or more hardware processors, the source document including numerical data and text, the source document displayed to a user in a graphical user interface (GUI) of a client device;   generating, using the one or more hardware processors, discovered data that relates to the numerical data included in the source document, the generating the discovered data comprising generating the discovered data based on at least a machine learning model trained on a corpus that includes articles in a domain related to the source document;   accessing, using the one or more hardware processors, a database trained with a machine learning algorithm to identify time series related data associated with the text; and   communicating the discovered data with a discovered data identifier, and the time series related data to the client device, via the communication network, for display in the GUI of the client device, wherein the display of the discovered data with the discovered data identifier, and the time series related data is displayed simultaneously with at least a portion of the source document in the GUI of the client device.

Join the waitlist — get patent alerts

Track US2019034835A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.