Approximate modeling of next combined result for stopping text-field recognition in a video stream
Abstract
Approximate modeling of next combined result for stopping text-field recognition in a video stream. In an embodiment, text-recognition results are generated from frames in a video stream and combined into an accumulated text-recognition result. A distance between the accumulated text-recognition result and a next accumulated text-recognition result is estimated based on an approximate model of the next accumulated text-recognition result, and a determination is made of whether or not to stop processing based on this estimated distance. After processing is stopped, the final accumulated text-recognition result may be output.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising using at least one hardware processor to:
until a determination to stop processing is made, for each of a plurality of image frames in a video stream,
receive the image frame,
generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters,
combine the text-recognition result with an accumulated text-recognition result,
estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and
determine whether or not to stop the processing based on the estimated distance; and,
after stopping the processing, output a character string based on the accumulated text-recognition result.
2 . The method of claim 1 , wherein estimating the distance between the accumulated text-recognition result and the future next accumulated text-recognition result comprises modeling the future next accumulated text-recognition result by using previous text-recognition results as candidates for the future next text-recognition result.
3 . The method of claim 2 , wherein estimating the distance between the accumulated text-recognition result and the future next accumulated text-recognition result further comprises, for each of the previous text-recognition results, calculating a distance between the accumulated text-recognition result and a combination of the accumulated text-recognition result with the previous text-recognition result.
4 . The method of claim 3 , wherein calculating a distance between the accumulated text-recognition result and the combination of the accumulated text-recognition result with the previous text-recognition result comprises aligning the accumulated text-recognition result with the previous text-recognition result based on a previous alignment of the accumulated text-recognition result with the previous text-recognition result.
5 . The method of claim 1 , wherein the at least one hardware processor is comprised in a mobile device, and wherein the image frames are received in real time or near-real time as the image frames are captured by a camera of the mobile device.
6 . A system comprising:
at least one hardware processor; and one or more software modules that, when executed by the at least one hardware processor,
until a determination to stop processing is made, for each of a plurality of image frames in a video,
receive the image frame,
generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters,
combine the text-recognition result with an accumulated text-recognition result,
estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and
determine whether or not to stop the processing based on the estimated distance, and,
after stopping the processing, output a character string based on the accumulated text-recognition result.
7 . A non-transitory computer-readable medium having instructions stored therein, wherein the instructions, when executed by a processor, cause the processor to:
until a determination to stop processing is made, for each of a plurality of image frames in a video,
receive the image frame,
generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters,
combine the text-recognition result with an accumulated text-recognition result,
estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and
determine whether or not to stop the processing based on the estimated distance; and,
after stopping the processing, output a character string based on the accumulated text-recognition result.Join the waitlist — get patent alerts
Track US2024029465A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.