US2013275612A1PendingUtilityA1

Systems and methods for scalable structured data distribution

Assignee: GOLDMAN SACHS & COPriority: Apr 13, 2012Filed: Apr 15, 2013Published: Oct 17, 2013
Est. expiryApr 13, 2032(~5.7 yrs left)· nominal 20-yr term from priority
G06F 16/24568G06F 16/24561G06F 16/25G06F 16/113H04L 65/60H04L 65/80G06F 16/2365H04L 65/75H04L 67/1001H04L 67/565H04L 67/561
51
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for efficiently absorbing, archiving, and distributing any size data sets are provided. Some embodiments provide flexible, policy-based distribution of high volume data through real time streaming as well as past data replay. In addition, some embodiments provide for a foundation of solid and unambiguous consistency across any vendor system through advanced version features. This consistency is particularly valuable to the financial industry, but also extremely useful to any company that manages multiple data distribution points for improved and reliable data availability.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 receiving streaming data from a data producer;   determining business aligned archive sequence from the streaming data that should be bundled together in accordance with a set of bundling parameters;   bundling the data into packages of data having a standard format;   ordering each of the packages of data using a series of consecutive integers produced by a master clock;   publishing metadata regarding availability of the packages of data on a control channel; and   delivering the packages of data to data consumers which have subscribed to the data producer.   
     
     
         2 . The method of  claim 1 , wherein the bundling parameters include declarative rules specified by a business. 
     
     
         3 . The method of  claim 1 , wherein the standard format is a standard neutral format biased for movement and compression of the streaming data. 
     
     
         4 . The method of  claim 1 , further comprising replaying the packages of data based on the ordering upon a request from a data consumer. 
     
     
         5 . The method of  claim 1 , further comprising archiving the packages of data in a platform independent manner. 
     
     
         6 . The method of  claim 1 , wherein the metadata that is published on the control channel includes indexes. 
     
     
         7 . The method of  claim 1 , wherein bundling the data includes identifying data that when compressed will result in each package in the packages of data having a desired size. 
     
     
         8 . The method of  claim 1 , wherein bundling the data includes associating new metadata with each of the bundled packages, wherein the new meta data comprises at least one of summary data, quality data, index data, or checksum data. 
     
     
         9 . The method of  claim 1 , wherein the delivery of the packages of data to the consumers comprises parallel delivery. 
     
     
         10 . The method of  claim 1 , further comprising using columnar checksums for verifying the data. 
     
     
         11 . The method of  claim 10 , wherein the columnar checksums allow for rounding errors with a specified tolerance. 
     
     
         12 . A system comprising:
 a bundler configured to receive streaming raw data from a data producer and bundle the raw data into a series of data packages and associate with each of the data packages a unique identifier having a monotonically increasing order based on upload from the data producer;   a transformer to receive the data packages having the associated unique identifier and generate loadable data structures for a reporting store associated with a data subscriber; and   a loader to receive and store the loadable data structures into a storage device associated with the data subscriber based on the monotonically increasing order.   
     
     
         13 . The system of  claim 12 , wherein the streaming raw data comprises multiple streams. 
     
     
         14 . The system of  claim 13 , wherein the data packages from each of the multiple streams are assigned different sets of unique identifiers. 
     
     
         15 . The system of  claim 13 , wherein each of the multiple streams of streaming raw data are assigned a flow priority. 
     
     
         16 . The system of  claim 12 , further comprising an identification module to receive a logical series of integers from the stream clock and generate the unique identifier having the logical ordering. 
     
     
         17 . The system of  claim 16 , further comprising a stream clock configured to generate the logical series of integers, and wherein a single integer is the unique identifier associated with a single data package in the series of data packages. 
     
     
         18 . The system of  claim 12 , further comprising:
 a data channel allowing data from a data producer to be continuously streamed to the data subscriber through the bundler;   a messaging channel to provide a current status of the data being continuously streamed from the data producer to the data subscriber; and   a control channel separate from the data channel to allow the data subscriber to request replay of the data.   
     
     
         19 . The system of  claim 18 , wherein the control channel is running at a faster rate than the data channel. 
     
     
         20 . The system of  claim 18 , wherein the control channel recursively publishes metadata regarding the data packages. 
     
     
         21 . The system of  claim 12 , further comprising an archiving service to archive the data packages. 
     
     
         22 . A method comprising:
 receiving a request to replay data bundled into data packages having a logical ordering assigned to the data packages before being stored in an archive, wherein the request includes a logical bound on the data to be replayed and identifies a format for the data subscriber;   retrieving, from the archive, data consistent with the logical bound; and   transforming the data packages into the loadable format identified in the request to replay the data.   
     
     
         23 . The method of  claim 22 , wherein data packages after the logical bound are ignored. 
     
     
         24 . The method of  claim 22 , further comprising:
 receiving a selection of an archiving strategy; and   archiving the data in accordance with the archiving strategy.   
     
     
         25 . The method of  claim 24 , wherein the archiving strategy is a dimension-based archiving strategy. 
     
     
         26 . The method of  claim 22 , wherein the data packages are compressed using columnar compression. 
     
     
         27 . The method of  claim 22 , wherein the data packages include metadata with each providing summary data, quality data, index data, or checksum data.

Join the waitlist — get patent alerts

Track US2013275612A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.