Computation platform agnostic data classification workflows
Abstract
Various embodiments include a classification platform system. A user can define a classification experiment on the classification platform system. For example, the user can define an input data space by selecting at least one of data sources interfaced with the classification platform system and defining a workflow configuration including a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow. The DG can specify how one or more outputs of each of the transformation blocks are fed into one or more other transformation blocks. The DG can be executed by various types of computation platforms. The classification platform system can schedule the experiment workflow to be executed on a distributed computation platform according to the input data space and the workflow configuration.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method, comprising:
defining a classification experiment for executing tasks on a computation platform by at least:
defining an input data space by selecting at least one of data sources interfaced with a classification platform system; and
defining, via a definition user interface of the classification platform system, a workflow configuration of the classification experiment by defining a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein the DG specifies one or more connections from one or more outputs of each transformation block in the experiment workflow to one or more other transformation blocks;
determining, based one or more previously executed classification experiments, that at least one transformation block included in the plurality of transformation blocks was previously executed; generating an execution schedule that specifies which system component executes each transformation block in the plurality of transformation blocks excluding the at least one transformation block; and scheduling the computation platform to execute the classification experiment by the system components under the execution schedule according to the input data space and the workflow configuration.
2 . The computer-implemented method of claim 1 , wherein the DG is arranged graphically via the definition user interface and wherein the definition user interface is a graphical user interface.
3 . The computer-implemented method of claim 1 , wherein said scheduling includes scheduling a first part of the workflow configuration to execute on a first computation platform and a second part of the workflow configuration to execute on a second computation platform.
4 . The computer-implemented method of claim 1 , further comprising selecting the computation platform based on a geographical or network location of the computation platform relative to one or more geographical or network locations of input data specified in the input data space.
5 . The computer-implemented method of claim 1 , further comprising maintaining a memorization database; wherein said scheduling includes preventing a transformation block from being executed by the computation platform when the transformation block as defined by the workflow configuration matches an entry in the memorization database; and wherein the entry includes pre-computed output result of the transformation block given the same input and configuration.
6 . The computer-implemented method of claim 1 , wherein that at least one data source comprises a live data source from a social networking system, and wherein the live data source produces an open-ended stream of new data entries formatted according to one or more data formats of the input data space.
7 . The computer-implemented method of claim 1 , wherein that at least one data source comprises a static data source from a social networking system, and wherein the static data source includes a static data set with a constant data size formatted according to one or more data formats of the input data space.
8 . The computer-implemented method of claim 1 , wherein the DG is acyclical and thereby prevents execution of the classification experiment to enter an infinite loop.
9 . The computer-implemented method of claim 1 , wherein a transformation block in the DG includes logic to dynamically modify the DG during execution of the experiment workflow.
10 . The computer-implemented method of claim 9 , wherein the transformation block includes logic to dynamically modify input data of an existing transformation block in the DG.
11 . The computer-implemented method of claim 9 , wherein the transformation block includes logic to change or remove an existing transformation block in the DG or to add a new transformation block to the DG.
12 . The computer-implemented method of claim 1 , further comprising piping an output result of executing the classification experiment to a social networking system to re-configure at least an application service of the social network system.
13 . The computer-implemented method of claim 1 , wherein the input data space is a labeled data space that includes at least a parameter to locate labeled data for training a supervised classifier machine learning model or for evaluating classification precision or recall of a classifier model, wherein said training or said evaluating is represented in a transformation block in the DG.
14 . The computer-implemented method of claim 1 , wherein the input data space is a prediction space that includes at least a parameter to locate input data to be classified in the classification experiment.
15 . The computer-implemented method of claim 1 , wherein defining the classification experiment further includes defining a domain configuration that includes at least a parameter binding the input data space to the workflow configuration.
16 . The computer-implemented method of claim 1 , wherein defining the classification experiment includes inheriting a directed graph for the workflow configuration from a workflow repository.
17 . One or more non-transitory computer readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:
defining a classification experiment for executing tasks on a computation platform by at least:
defining an input data space by selecting at least one of data sources interfaced with a classification platform system; and
defining, via a definition user interface of the classification platform system, a workflow configuration of the classification experiment by defining a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein the DG specifies one or more connections from one or more outputs of each transformation block in the experiment workflow to one or more other transformation blocks;
determining, based one or more previously executed classification experiments, that at least one transformation block included in the plurality of transformation blocks was previously executed; generating an execution schedule that specifies which system component executes each transformation block in the plurality of transformation blocks excluding the at least one transformation block; and scheduling the computation platform to execute the classification experiment by the system components under the execution schedule according to the input data space and the workflow configuration.
18 . The one or more non-transitory computer readable media of claim 17 , wherein the instructions further cause the one or more processors to maintain a memorization database; wherein said scheduling includes preventing a transformation block from being executed by the computation platform when the transformation block as defined by the workflow configuration matches an entry in the memorization database; and wherein the entry includes pre-computed output result of the transformation block given the same input and configuration.
19 . A computer system, comprising:
one or more memories storing one or more instructions; and one or more processors for executing the one or more instructions to: define a classification experiment for executing tasks on a computation platform by at least:
defining an input data space by selecting at least one of data sources interfaced with a classification platform system; and
defining, via a definition user interface of the classification platform system, a workflow configuration of the classification experiment by defining a directed graph (DG) connecting a plurality of transformation blocks to represent an experiment workflow, wherein the DG specifies one or more connections from one or more outputs of each transformation block in the experiment workflow to one or more other transformation blocks;
determine, based one or more previously executed classification experiments, that at least one transformation block included in the plurality of transformation blocks was previously executed; generate an execution schedule that specifies which system component executes each transformation block in the plurality of transformation blocks excluding the at least one transformation block; and schedule the computation platform to execute the classification experiment by the system components under the execution schedule according to the input data space and the workflow configuration.
20 . The computer system of claim 19 , wherein the one or more processors further maintain a memorization database; wherein said scheduling includes preventing a transformation block from being executed by the computation platform when the transformation block as defined by the workflow configuration matches an entry in the memorization database; and wherein the entry includes pre-computed output result of the transformation block given the same input and configuration.Join the waitlist — get patent alerts
Track US2020334293A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.