Recognition of handwritten text via neural networks
Abstract
In one embodiment, a system receives an image depicting a line of text. The system segments the image into two or more fragment images. For each of the two or more fragment images, the system determines a first hypothesis to segment the fragment image into a first plurality of grapheme images and a first fragmentation confidence score. The system determines a second hypothesis to segment the fragment image into a second plurality of grapheme images and a second fragmentation confidence score. The system determines that the first fragmentation confidence score is greater than the second fragmentation confidence score. The system translates the first plurality of grapheme images defined by the first hypothesis to symbols. The system assembles the symbols of each fragment image to derive the line of text.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
generating one or more hypotheses, each of the one or more hypotheses segmenting an image of a text into a plurality of fragment images, wherein a fragment image of the plurality of fragment images depicts one or more words of the text; obtaining one or more fragmentation confidence scores, each fragmentation confidence score obtained for a respective hypothesis of the one or more hypotheses, by:
applying a recognition model to the respective plurality of fragment images to identify (i) a plurality of symbols corresponding to the respective plurality of fragment images, and (ii) a plurality of classification confidence scores associated with the respective plurality of fragment images; and
determining, using the plurality of classification confidence scores, the fragmentation confidence score for the respective hypothesis; and
using the one or more fragmentation confidence scores, selecting the plurality of symbols, identified for a winning hypothesis of the one or more hypotheses, as a recognized text.
2 . The method of claim 1 , further comprising:
determining, using a language detection model, a language associated with the image; and selecting the plurality of symbols corresponding to the respective plurality of fragment images from a corpus of symbols of the determined language.
3 . The method of claim 1 , further comprising:
for each of the one or more hypotheses:
applying a structural classification model to the respective plurality of fragment images to identify an additional plurality of classification confidence scores characterizing structural similarity of a respective fragment image to one or more reference images; and
wherein the fragmentation confidence score for the respective hypothesis is further determined using the additional plurality of classification confidence scores.
4 . The method of claim 1 , wherein the recognition model is trained using a loss function, wherein the loss function comprises one or more of:
a cross entropy loss function, a center loss function, or a close-to-center penalty loss function.
5 . The method of claim 1 , wherein an input into the recognition model comprises:
a first input comprising one or more fragment images of the respective plurality of fragment images, and a second input comprising one or more geometric features of the one or more fragment images.
6 . The method of claim 5 , wherein the one or more geometric features comprise at least one aspect ratio for one or more graphemes in the one or more fragment images.
7 . The method of claim 1 , further comprising:
validating the plurality of symbols, identified for the winning hypothesis, using one or more of a morphological model, a dictionary model, or a syntactical model.
8 . A system, comprising:
a memory; a processing device, communicatively coupled to the memory, the processing device to:
generate one or more hypotheses, each of the one or more hypotheses segmenting an image of a handwritten text into a plurality of fragment images, wherein a fragment image of the plurality of fragment images depicts one or more words of the text;
obtain one or more fragmentation confidence scores, each fragmentation confidence score obtained for a respective hypothesis of the one or more hypotheses, by:
applying a recognition model to the respective plurality of fragment images to identify (i) a plurality of symbols corresponding to the respective plurality of fragment images, and (ii) a plurality of classification confidence scores associated with the respective plurality of fragment images; and
determining, using the plurality of classification confidence scores, the fragmentation confidence score for the respective hypothesis; and
select, using the one or more fragmentation confidence scores, the plurality of symbols, identified for a winning hypothesis of the one or more hypotheses, as a recognized text.
9 . The system of claim 8 , wherein the processing device is further to:
determine, using a language detection model, a language associated with the image; and select the plurality of symbols corresponding to the respective plurality of fragment images from a corpus of symbols of the determined language.
10 . The system of claim 8 , wherein the processing device is further to:
for each of the one or more hypotheses:
apply a structural classification model to the respective plurality of fragment images to identify an additional plurality of classification confidence scores characterizing structural similarity of a respective fragment image to one or more reference images; and
wherein the fragmentation confidence score for the respective hypothesis is further determined using the additional plurality of classification confidence scores.
11 . The system of claim 8 , wherein the recognition model is trained using a loss function, wherein the loss function comprises one or more of:
a cross entropy loss function, a center loss function, or a close-to-center penalty loss function.
12 . The system of claim 8 , wherein an input into the recognition model comprises:
a first input comprising one or more fragment images of the respective plurality of fragment images, and a second input comprising one or more geometric features of the one or more fragment images.
13 . The system of claim 12 , wherein the one or more geometric features comprise at least one aspect ratio for one or more graphemes in the one or more fragment images.
14 . The system of claim 8 , wherein the processing device is further to:
validate the plurality of symbols, identified for the winning hypothesis, using one or more of a morphological model, a dictionary model, or a syntactical model.
15 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processing device system, cause the processing device to:
generate one or more hypotheses, each of the one or more hypotheses segmenting an image of a handwritten text into a plurality of fragment images, wherein a fragment image of the plurality of fragment images depicts one or more words of the text; obtain one or more fragmentation confidence scores, each fragmentation confidence score obtained for a respective hypothesis of the one or more hypotheses, by:
applying a recognition model to the respective plurality of fragment images to identify (i) a plurality of symbols corresponding to the respective plurality of fragment images, and (ii) a plurality of classification confidence scores associated with the respective plurality of fragment images; and
determining, using the plurality of classification confidence scores, the fragmentation confidence score for the respective hypothesis; and
select, using the one or more fragmentation confidence scores, the plurality of symbols, identified for a winning hypothesis of the one or more hypotheses, as a recognized text.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions are further to cause the processing device to:
determine, using a language detection model, a language associated with the image; and select the plurality of symbols corresponding to the respective plurality of fragment images from a corpus of symbols of the determined language.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions are further to cause the processing device to:
for each of the one or more hypotheses:
apply a structural classification model to the respective plurality of fragment images to identify an additional plurality of classification confidence scores characterizing structural similarity of a respective fragment image to one or more reference images; and
wherein the fragmentation confidence score for the respective hypothesis is further determined using the additional plurality of classification confidence scores.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein an input into the recognition model comprises:
a first input comprising one or more fragment images of the respective plurality of fragment images, and a second input comprising one or more geometric features of the one or more fragment images.
19 . The non-transitory computer-readable storage medium of claim 18 , wherein the one or more geometric features comprise at least one aspect ratio for one or more graphemes in the one or more fragment images.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions are further to cause the processing device to:
validate the plurality of symbols, identified for the winning hypothesis, using one or more of a morphological model, a dictionary model, or a syntactical model.Join the waitlist — get patent alerts
Track US2024037969A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.