Mining procedure dialogs from source content
Abstract
An embodiment of the invention may include a method, computer program product and computer system for human-machine communication. The method, computer program product and computer system may include a computing device that maps linguistic data of source content to a vector. The computing device may cluster the linguistic data of source content. The computing device may determine a plurality of segments based on the mapped linguistic data and the clustered linguistic data. The computing device may transform a segment of the plurality of segments into representative data, the representative data is a function of the remaining plurality of segments.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for human-machine communication, the method comprising:
mapping linguistic data of the source content to a vector; clustering the linguistic data of the source content; determining a plurality of segments based on the mapped linguistic data and the clustered linguistic data; and transforming a segment of the plurality of segments into representative data, the representative data is a function of the remaining plurality of segments.
2 . The method of claim 1 , further comprising:
identifying related segments of linguistic data based on cohesions; generating linguistic data blocks comprised of the related segments of linguistic data; identifying similarities between the linguistic data blocks based on cohesions; separating the linguistic data blocks into groups based on the cohesions identified; generating a procedure dialogue based on the cohesions between data blocks of the same group; receiving user input; identifying the topic of the user input; identifying a procedure dialogue based on the topic of the user input; and responding to the user input based on the procedure dialogue identified.
3 . The method of claim 2 , further comprising:
identifying cohesions between different procedure dialogues based on proximity within the source content.
4 . The method of claim 1 , wherein the source content may comprise audio, visual and textual content.
5 . The method of claim 1 , wherein the linguistic data is mapped to the vector in segments.
6 . The method of claim 4 , wherein the segments of linguistic data may comprise sentence segments.
7 . The method of claim 4 , wherein the segments of linguistic data may comprise word segments.
8 . The method of claim 1 , wherein the segments of linguistic data are mapped using a neural network.
9 . The method of claim 8 , wherein the neural network is contextual continuous bag-of-words architecture.
10 . The method of claim 8 , wherein the neural network is contextual skip-gram architecture.
11 . The method of claim 2 , wherein the cohesions are linguistic similarities measured by manually set thresholds.
12 . The method of claim 2 , wherein the procedure dialogue comprises nodes and flows between the nodes, wherein the nodes comprise questions to be presented to the user and the flows represent possible user answers.
13 . A computer program product for human-machine communication, the computer program product comprising:
a computer-readable storage device and program instructions stored on computer-readable storage device, the program instructions comprising:
program instructions to map linguistic data of the source content to a vector;
program instructions to cluster the linguistic data of the source content;
program instructions to determine a plurality of segments based on the mapped linguistic data and the clustered linguistic data; and
program instructions to transform a segment of the plurality of segments into representative data, the representative data is a function of the remaining plurality of segments.
14 . The computer program product of claim 13 , further comprising:
program instructions to identify related segments of linguistic data based on cohesions; program instructions to generate linguistic data blocks comprised of the related segments of linguistic data; program instructions to identify similarities between the linguistic data blocks based on cohesions; program instructions to separate the linguistic data blocks into groups based on the cohesions identified; program instructions to generate a procedure dialogue based on the cohesions between data blocks of the same group; program instructions to receive user input; program instructions to identify the topic of the user input; program instructions to identify a procedure dialogue based on the topic of the user input; and program instructions to respond to the user input based on the procedure dialogue identified.
15 . The computer program product of claim 14 , further comprising:
program instructions to identify cohesions between different procedure dialogues based on proximity within the source content.
16 . The computer program product of claim 13 , wherein the segments of linguistic data are mapped using a neural network.
17 . A computer system for human-machine communication, the computer system comprising:
one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, the program instructions comprising:
program instructions to map linguistic data of the source content to a vector;
program instructions to cluster the linguistic data of the source content;
program instructions to determine a plurality of segments based on the mapped linguistic data and the clustered linguistic data; and
program instructions to transform a segment of the plurality of segments into representative data, the representative data is a function of the remaining plurality of segments.
18 . The computer system claim 17 , further comprising:
program instructions to identify related segments of linguistic data based on cohesions; program instructions to generate linguistic data blocks comprised of the related segments of linguistic data; program instructions to identify similarities between the linguistic data blocks based on cohesions; program instructions to separate the linguistic data blocks into groups based on the cohesions identified; program instructions to generate a procedure dialogue based on the cohesions between data blocks of the same group; program instructions to receive user input; program instructions to identify the topic of the user input; program instructions to identify a procedure dialogue based on the topic of the user input; and program instructions to respond to the user input based on the procedure dialogue identified.
19 . The computer system of claim 18 further comprising:
program instructions to identify cohesions between different procedure dialogues based on proximity within the source content.
20 . The computer system of claim 17 , wherein the segments of linguistic data are mapped using a neural network.Join the waitlist — get patent alerts
Track US2019026346A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.