US2024046916A1PendingUtilityA1
Graphics translation to natural language
Assignee: TOSHIBA GLOBAL COMMERCE SOLUTIONS HOLDINGS CORPPriority: Jul 27, 2021Filed: Oct 23, 2023Published: Feb 8, 2024
Est. expiryJul 27, 2041(~15 yrs left)· nominal 20-yr term from priority
G10L 13/027G06V 10/10G06F 40/30G10L 13/00G06Q 30/0601G06F 40/279G06Q 30/0282G06Q 30/015G06T 11/00G06F 21/44G06F 18/2415G06Q 30/0631G06Q 10/101
54
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
The present disclosure provides techniques for graphics translation. A subset of a plurality of natural language image descriptions for an image of a product is received. A set of shared natural language descriptors is identified in the subset of the plurality of natural language image descriptions. The set of shared natural language descriptors is aggregated, and a description for the first image is generated based on the aggregated set of shared natural language image descriptions.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
receiving, for a first image of a first product, a subset of a plurality of natural language image descriptions; identifying a set of shared natural language descriptors in the subset of the plurality of natural language image descriptions; aggregating the set of shared natural language descriptors; and generating a description for the first image based on the aggregated set of shared natural language image descriptions.
2 . The method of claim 1 , wherein the plurality of natural language image descriptions correspond to user-provided reviews of the first product.
3 . The method of claim 1 , wherein the plurality of natural language image descriptions are collected by:
providing the first image to a plurality of users; and providing a form to the plurality of users, wherein the form requests user-input on specified descriptors for the first image.
4 . The method of claim 1 , further comprising:
determining a first retailer-provided description of the first product; determining a second retailer-provided description of a second product depicted in a second image; and upon determining that at least a portion of the first and second retailer-provided descriptions match, using at least a portion of the set of shared natural language descriptors as a description for the second product.
5 . The method of claim 1 , further comprising training one or more machine learning models to generate descriptions for images, based at least in part on the first image and the set of shared natural language descriptors.
6 . The method of claim 1 , further comprising:
receiving a first request to provide a description of the first image; and returning the description in response to the first request.
7 . The method of claim 1 , wherein identifying the set of shared natural language descriptors comprises at least one of:
(i) determining that a natural language descriptor is present in at least a threshold number or threshold percentage of the plurality of natural language image descriptions, or (ii) determining that a natural language descriptor is a most common natural language descriptor in the plurality of natural language image descriptions.
8 . A computer-readable storage medium containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising:
receiving, for a first image of a first product, a subset of a plurality of natural language image descriptions; identifying a set of shared natural language descriptors in the subset of the plurality of natural language image descriptions; aggregating the set of shared natural language descriptors; and generating a description for the first image based on the aggregated set of shared natural language image descriptions.
9 . The computer-readable storage medium of claim 8 , wherein the plurality of natural language image descriptions correspond to user-provided reviews of the first product.
10 . The computer-readable storage medium of claim 8 , wherein the plurality of natural language image descriptions are collected by:
providing the first image to a plurality of users; and providing a form to the plurality of users, wherein the form requests user-input on specified descriptors for the first image.
11 . The computer-readable storage medium of claim 8 , the operation further comprising:
determining a first retailer-provided description of the first product; determining a second retailer-provided description of a second product depicted in a second image; and upon determining that at least a portion of the first and second retailer-provided descriptions match, using at least a portion of the set of shared natural language descriptors as a description for the second product.
12 . The computer-readable storage medium of claim 8 , the operation further comprising training one or more machine learning models to generate descriptions for images, based at least in part on the first image and the set of shared natural language descriptors.
13 . The computer-readable storage medium of claim 8 , the operation further comprising:
receiving a first request to provide a description of the first image; and returning the description in response to the first request.
14 . The computer-readable storage medium of claim 8 , wherein identifying the set of shared natural language descriptors comprises at least one of:
(i) determining that a natural language descriptor is present in at least a threshold number or threshold percentage of the plurality of natural language image descriptions, or (ii) determining that a natural language descriptor is a most common natural language descriptor in the plurality of natural language image descriptions.
15 . A system comprising:
one or more computer processors; and a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising:
receiving, for a first image of a first product, a subset of a plurality of natural language image descriptions;
identifying a set of shared natural language descriptors in the subset of the plurality of natural language image descriptions;
aggregating the set of shared natural language descriptors; and
generating a description for the first image based on the aggregated set of shared natural language image descriptions.
16 . The system of claim 15 , wherein the plurality of natural language image descriptions correspond to user-provided reviews of the first product.
17 . The system of claim 15 , wherein the plurality of natural language image descriptions are collected by:
providing the first image to a plurality of users; and providing a form to the plurality of users, wherein the form requests user-input on specified descriptors for the first image.
18 . The system of claim 15 , the operation further comprising:
determining a first retailer-provided description of the first product; determining a second retailer-provided description of a second product depicted in a second image; and upon determining that at least a portion of the first and second retailer-provided descriptions match, using at least a portion of the set of shared natural language descriptors as a description for the second product.
19 . The system of claim 15 , the operation further comprising training one or more machine learning models to generate descriptions for images, based at least in part on the first image and the set of shared natural language descriptors.
20 . The system of claim 15 , wherein identifying the set of shared natural language descriptors comprises at least one of:
(i) determining that a natural language descriptor is present in at least a threshold number or threshold percentage of the plurality of natural language image descriptions, or (ii) determining that a natural language descriptor is a most common natural language descriptor in the plurality of natural language image descriptions.Join the waitlist — get patent alerts
Track US2024046916A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.