US2019354554A1PendingUtilityA1

Graphically managing data classification workflows in a social networking system with directed graphs

Assignee: FACEBOOK INCPriority: Jun 30, 2016Filed: Aug 2, 2019Published: Nov 21, 2019
Est. expiryJun 30, 2036(~9.9 yrs left)· nominal 20-yr term from priority
G06Q 10/40G06F 16/285G06Q 10/04G06F 16/9024G06Q 50/01
64
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In one embodiment, a method includes defining an input data space, defining a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein at least one of the plurality of transformation blocks includes logic to dynamically modify the DG during execution of the experiment workflow, formatting the DG and the input data space into a data structure such that the data structure is interpretable by a plurality of different computation platforms, scheduling a distributed computation platform selected from the plurality of different computation platforms to execute the experiment workflow according to the input data space and the DG, and imperatively programming computing nodes of the distributed computation platform to execute the experiment workflow based on the DG if at least one of the transformation blocks dynamically modifies the DG during the execution of the experiment workflow.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 defining an input data space;   defining a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein at least one of the plurality of transformation blocks includes logic to dynamically modify the DG during execution of the experiment workflow;   formatting the DG and the input data space into a data structure such that the data structure is interpretable by a plurality of different computation platforms;   scheduling a distributed computation platform selected from the plurality of different computation platforms to execute the experiment workflow according to the input data space and the DG; and   imperatively programming computing nodes of the distributed computation platform to execute the experiment workflow based on the DG if at least one of the transformation blocks dynamically modifies the DG during the execution of the experiment workflow.   
     
     
         2 . The method of  claim 1 , wherein defining the input data space comprises:
 interfacing one or more data sources in a communication system to a classification platform system; and   selecting at least one of the data sources interfaced with the classification platform system.   
     
     
         3 . The method of  claim 2 , wherein the data sources comprise one or more live data sources from the communication system, and wherein the live data sources produce an open-ended stream of new data entries formatted according to one or more data formats of the defined input data space. 
     
     
         4 . The method of  claim 2 , wherein the data sources comprise one or more static data sources from the communication system, and wherein the static data sources comprise a static data set with a constant data size formatted according to one or more data formats of the defined input data space. 
     
     
         5 . The method of  claim 2 , wherein the input data space selects at least a live data source from the data sources to feed into at least one of the transformation blocks, and wherein the distributed computation platform is configured to execute the experiment workflow in real-time in response to new data from the live data source. 
     
     
         6 . The method of  claim 1 , wherein scheduling the distributed computation platform to execute the experiment workflow comprises scheduling a first part of the experiment workflow to execute on a first computation platform and a second part of the experiment workflow to execute on a second computation platform. 
     
     
         7 . The method of  claim 1 , wherein the distributed computation platform is selected based on a geographical or network location of the distributed computation platform relative to one or more geographical or network locations of input data specified in the input data space. 
     
     
         8 . The method of  claim 1 , further comprising:
 maintaining a memorization database, wherein scheduling the distributed computation platform to execute the experiment workflow comprises preventing a transformation block from being executed by the distributed computation platform when the transformation block as defined by the experiment workflow matches an entry in the memorization database, and wherein the entry comprises pre-computed output result of the transformation block given the same input and experiment workflow.   
     
     
         9 . The method of  claim 1 , wherein the DG is acyclical and thereby prevents execution of the experiment workflow to enter an infinite loop. 
     
     
         10 . The method of  claim 1 , wherein a transformation block in the DG comprises logic to dynamically modify the DG during execution of the experiment workflow. 
     
     
         11 . The method of  claim 10 , wherein the transformation block comprises logic to dynamically modify input data of an existing transformation block in the DG. 
     
     
         12 . The method of  claim 10 , wherein the transformation block comprises logic to change or remove an existing transformation block in the DG or to add a new transformation block to the DG. 
     
     
         13 . The method of  claim 1 , further comprising:
 piping an output result of executing the experiment workflow to a communication system to reconfigure at least an application service of the communication system.   
     
     
         14 . The method of  claim 1 , wherein the input data space is a labeled data space that comprises at least a parameter to locate labeled data for training a supervised classifier or for evaluating classification precision or recall of a classifier, wherein the training or evaluating is represented in a transformation block in the DG. 
     
     
         15 . The method of  claim 1 , wherein the input data space is a prediction space that comprises at least a parameter to locate input data to be classified in the experiment workflow. 
     
     
         16 . The method of  claim 1 , further comprising:
 defining a domain configuration that comprises at least a parameter binding the input data space to the experiment workflow.   
     
     
         17 . The method of  claim 1 , further comprising:
 inheriting the DG for the experiment workflow from a workflow repository.   
     
     
         18 . The method of  claim 1 , wherein the transformation blocks comprise one or more of a data feature extraction process, a data feature filtering process, a data feature transformation process, a classifier deliberation process, a classifier training process, or a classifier evaluation process. 
     
     
         19 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
 define an input data space;   define a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein at least one of the plurality of transformation blocks includes logic to dynamically modify the DG during execution of the experiment workflow;   format the DG and the input data space into a data structure such that the data structure is interpretable by a plurality of different computation platforms;   schedule a distributed computation platform selected from the plurality of different computation platforms to execute the experiment workflow according to the input data space and the DG; and   imperatively program computing nodes of the distributed computation platform to execute the experiment workflow based on the DG if at least one of the transformation blocks dynamically modifies the DG during the execution of the experiment workflow.   
     
     
         20 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
 define an input data space;   define a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein at least one of the plurality of transformation blocks includes logic to dynamically modify the DG during execution of the experiment workflow;   format the DG and the input data space into a data structure such that the data structure is interpretable by a plurality of different computation platforms;   schedule a distributed computation platform selected from the plurality of different computation platforms to execute the experiment workflow according to the input data space and the DG; and   imperatively program computing nodes of the distributed computation platform to execute the experiment workflow based on the DG if at least one of the transformation blocks dynamically modifies the DG during the execution of the experiment workflow.

Join the waitlist — get patent alerts

Track US2019354554A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.