System and method for knowledge-based audio-text modeling via automatic multimodal graph construction
Abstract
Knowledge-based audio-text modeling via automatic multimodal graph construction is performed. An audio dataset is received, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip of the audio data. Graph nodes of interest are identified from a sematic network, the graph nodes being descriptive of semantics of the knowledge domain of the contents of the audio dataset. A large language model (LLM) is utilized for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph. The extracted knowledge graph is validated utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for knowledge-based audio-text modeling via automatic multimodal graph construction, comprising:
receiving an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of audio contents of the respective clip; identifying graph nodes of interest from a sematic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset; utilizing a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph; validating the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and utilizing the knowledge graph, as validated, for a downstream application.
2 . The method of claim 1 , wherein the metadata includes human annotations describing the audio contents of the respective clips of the audio data.
3 . The method of claim 1 , wherein the metadata includes machine-learned labels, attributes, and/or other forms of recognition outcomes inferred from audio and/or speech data using one or more machine learning models.
4 . The method of claim 1 , wherein the graph nodes of interest are one or more of: defined based on user knowledge of the knowledge domain, queried from the sematic network as graph nodes describing semantics of the knowledge domain; extracted from a database of domain knowledge; received from the LLM responsive to a prompt for relevant graph nodes for the knowledge domain.
5 . The method of claim 1 , wherein inferring the supplemental data includes receiving the supplemental data from the LLM responsive to a prompt for requesting the LLM to infer content for names of the graph nodes for which there is no metadata available.
6 . The method of claim 1 , wherein the downstream application includes an audio classification application using the knowledge graph for sound event detection and/or audio tagging.
7 . The method of claim 1 , wherein the downstream application includes an audio captioning application using the knowledge graph for audio retrieval.
8 . The method of claim 1 , wherein the downstream application includes representing the knowledge graph as an adjacency matrix to perform multimodal graph representation learning.
9 . The method of claim 1 , wherein the downstream application includes using the knowledge graph to define knowledge-based clusters for contrastive learning.
10 . The method of claim 1 , wherein the downstream application includes using the knowledge graph to curate controllable prompts, captions, and/or descriptive contents for building knowledge-guided generative models.
11 . A system for knowledge-based audio-text modeling via automatic multimodal graph construction, comprising:
one or more hardware computing devices configured to:
receive an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip;
identify graph nodes of interest from a sematic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset;
utilize a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph;
validate the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and
utilize the knowledge graph, as validated, for a downstream application.
12 . The system of claim 11 , wherein the metadata includes human annotations describing the audio contents of the respective clips of the audio data.
13 . The system of claim 11 , wherein the metadata includes machine-learned labels, attributes, and/or other forms of recognition outcomes inferred from audio and/or speech data using one or more machine learning models.
14 . The system of claim 11 , wherein the graph nodes of interest are one or more of: defined based on user knowledge of the knowledge domain, queried from the sematic network as graph nodes describing semantics of the knowledge domain; extracted from a database of domain knowledge; received from the LLM responsive to a prompt for relevant graph nodes for the knowledge domain.
15 . The system of claim 11 , wherein inferring the supplemental data includes receiving the supplemental data from the LLM responsive to a prompt for requesting the LLM to infer content for names of the graph nodes for which there is no metadata available.
16 . The system of claim 11 , wherein the downstream application includes an audio classification application using the knowledge graph for sound event detection and/or audio tagging.
17 . The system of claim 11 , wherein the downstream application includes an audio captioning application using the knowledge graph for audio retrieval.
18 . The system of claim 11 , wherein the downstream application includes representing the knowledge graph as an adjacency matrix to perform multimodal graph representation learning.
19 . The system of claim 11 , wherein the downstream application includes using the knowledge graph to define knowledge-based clusters for contrastive learning.
20 . The system of claim 11 , wherein the downstream application includes using the knowledge graph to curate controllable prompts/captions/descriptive contents for building knowledge-guided generative models.
21 . A non-transitory computer-readable medium comprising instructions for a knowledge-based audio-text modeling via automatic multimodal graph construction that, when executed by one or more hardware computing devices cause the one or more hardware computing devices to perform operations including to:
receive an audio dataset, the audio dataset including clips of audio data, wherein each of the clips of the audio data is paired with corresponding metadata descriptive of the audio contents of the respective clip; identify graph nodes of interest from a sematic network, the graph nodes being descriptive of semantics of a knowledge domain of the audio dataset; utilize a large language model (LLM) for categorizing the metadata into the graph nodes and for inferring supplemental data for the graph nodes for which there is no metadata, producing an extracted knowledge graph; validate the extracted knowledge graph utilizing the LLM to perform relation verification of edges between the graph nodes of the extracted knowledge graph, thereby mitigating hallucination effects in the categorizing and inferring of the supplemental data; and utilize the knowledge graph, as validated, for a downstream application.Join the waitlist — get patent alerts
Track US2025335705A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.