System and method for data organization, optimization and analytics
Abstract
A system and method for data organization, optimization and analytics includes a web server, thrift server, distributed processing framework, key value store, distributed file system, and relational database. The web server provides a method whereby users issue control actions and query for records via interaction with the thrift server. The thrift server is the center of coordination and communication for the system and interacts with other system elements. The key value store organizes all of the operational data for the system. The key value store runs on a highly scalable distributed system, including a distributed file system for storage of data on disk. The distributed processing framework enables data to be processed in bulk and is used to execute analytical processing on the data. The relational database hold all of the administrative data in the system. Search queries are submitted by end user and results of the search query are sent from the web server to the end user. The web server sends control actions to queue background map reduce jobs. These jobs run in the distributed processing framework and are used to write data and indexes and execute bulk analytics against the key value store.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method for efficiently importing and indexing a plurality of data sets, each data set of the plurality having a schema different from every other data set of the plurality, each data set of the plurality comprising at least one field containing at least one value, the at least one value having a type and length, the method comprising:
automatically identifying the physical format of each data set of the plurality; for each data set of the plurality, automatically identifying the at least one field and at least one value; for each data set of the plurality, automatically identifying a name of the at least one field; and storing each value indexed to the name of the field containing each said value in a data set of the plurality.
2 . The method of claim 1 , further comprising for each data set of the plurality, automatically identifying the type of the at least one value.
3 . The method of claim 2 , further comprising storing each type of the at least one value indexed to the name of the field containing each said value in a data set of the plurality.
4 . The method of claim 2 , wherein the type comprises a numerical string.
5 . The method of claim 1 , wherein each data set of the plurality is in a file format of a plurality of file formats, and the at least one field and at least one value of each data set of the plurality are automatically identified based on the file format of each said data set.
6 . The method of claim 2 , wherein each data set of the plurality is in a file format of a plurality of file formats, and the type of the at least one value of each data set of the plurality is automatically identified based on the file format of each said data set.
7 . The method of claim 1 , wherein each value is stored in a sorted key-value store.
8 . At least one computer-readable medium on which are stored instructions that, when executed by at least one processing device, enable the at least one processing device to perform a method, comprising the steps of:
automatically identifying the physical format of each data set of the plurality; for each data set of the plurality, automatically identifying the at least one field and at least one value; for each data set of the plurality, automatically identifying a name of the at least one field; and storing each value indexed to the name of the field containing each said value in a data set of the plurality.
9 . The method of claim 8 , further comprising for each data set of the plurality, automatically identifying the type of the at least one value.
10 . The method of claim 9 , further comprising storing each type of the at least one value indexed to the name of the field containing each said value in a data set of the plurality.
11 . The method of claim 9 , wherein the type comprises a numerical string.
12 . The method of claim 8 , wherein each data set of the plurality is in a file format of a plurality of file formats, and the at least one field and at least one value of each data set of the plurality are automatically identified based on the file format of each said data set.
13 . The method of claim 9 , wherein each data set of the plurality is in a file format of a plurality of file formats, and the type of the at least one value of each data set of the plurality is automatically identified based on the file format of each said data set.
14 . The method of claim 8 , wherein each value is stored in a sorted key-value store.Join the waitlist — get patent alerts
Track US2021056102A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.