US2024029465A1PendingUtilityA1

Approximate modeling of next combined result for stopping text-field recognition in a video stream

Assignee: Smart Engines Service LLCPriority: Jul 7, 2020Filed: Oct 5, 2023Published: Jan 25, 2024
Est. expiryJul 7, 2040(~13.9 yrs left)· nominal 20-yr term from priority
G06V 30/413G06F 16/9027G06V 10/20G06V 10/25G06V 20/62G06V 20/46
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Approximate modeling of next combined result for stopping text-field recognition in a video stream. In an embodiment, text-recognition results are generated from frames in a video stream and combined into an accumulated text-recognition result. A distance between the accumulated text-recognition result and a next accumulated text-recognition result is estimated based on an approximate model of the next accumulated text-recognition result, and a determination is made of whether or not to stop processing based on this estimated distance. After processing is stopped, the final accumulated text-recognition result may be output.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising using at least one hardware processor to:
 until a determination to stop processing is made, for each of a plurality of image frames in a video stream,
 receive the image frame, 
 generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters, 
 combine the text-recognition result with an accumulated text-recognition result, 
 estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and 
 determine whether or not to stop the processing based on the estimated distance; and, 
   after stopping the processing, output a character string based on the accumulated text-recognition result.   
     
     
         2 . The method of  claim 1 , wherein estimating the distance between the accumulated text-recognition result and the future next accumulated text-recognition result comprises modeling the future next accumulated text-recognition result by using previous text-recognition results as candidates for the future next text-recognition result. 
     
     
         3 . The method of  claim 2 , wherein estimating the distance between the accumulated text-recognition result and the future next accumulated text-recognition result further comprises, for each of the previous text-recognition results, calculating a distance between the accumulated text-recognition result and a combination of the accumulated text-recognition result with the previous text-recognition result. 
     
     
         4 . The method of  claim 3 , wherein calculating a distance between the accumulated text-recognition result and the combination of the accumulated text-recognition result with the previous text-recognition result comprises aligning the accumulated text-recognition result with the previous text-recognition result based on a previous alignment of the accumulated text-recognition result with the previous text-recognition result. 
     
     
         5 . The method of  claim 1 , wherein the at least one hardware processor is comprised in a mobile device, and wherein the image frames are received in real time or near-real time as the image frames are captured by a camera of the mobile device. 
     
     
         6 . A system comprising:
 at least one hardware processor; and   one or more software modules that, when executed by the at least one hardware processor,
 until a determination to stop processing is made, for each of a plurality of image frames in a video,
 receive the image frame, 
 generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters, 
 combine the text-recognition result with an accumulated text-recognition result, 
 estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and 
 determine whether or not to stop the processing based on the estimated distance, and, 
 
 after stopping the processing, output a character string based on the accumulated text-recognition result. 
   
     
     
         7 . A non-transitory computer-readable medium having instructions stored therein, wherein the instructions, when executed by a processor, cause the processor to:
 until a determination to stop processing is made, for each of a plurality of image frames in a video,
 receive the image frame, 
 generate a text-recognition result from the image frame, wherein the text-recognition result comprises a vector of class estimations for each of one or more characters, 
 combine the text-recognition result with an accumulated text-recognition result, 
 estimate a distance between the accumulated text-recognition result and a future next accumulated text-recognition result, which would be generated from an image frame that has not yet been received, based on an approximate model of the future next accumulated text-recognition result, and 
 determine whether or not to stop the processing based on the estimated distance; and, 
   after stopping the processing, output a character string based on the accumulated text-recognition result.

Join the waitlist — get patent alerts

Track US2024029465A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.