US2011091098A1PendingUtilityA1

System and Method for Detecting Text in Real-World Color Images

Assignee: YUILLE ALANPriority: Sep 2, 2005Filed: Oct 18, 2010Published: Apr 21, 2011
Est. expirySep 2, 2025(expired)· nominal 20-yr term from priority
G06V 30/10G06V 20/63G06V 20/62
41
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method and apparatus for detecting text in real-world images comprises calculating a cascade of classifiers, the cascade comprising a plurality of stages, each stage including one or more weak classifiers, the plurality of stages organized to start out with classifiers that are most useful for ruling out non-text regions, and removing regions classified as non-text regions from the cascade prior to completion of the cascade, to further speed up processing.

Claims

exact text as granted — not AI-modified
1 . A method of detecting text in real-world images comprising:
 dividing an image representing a real-world scene into one or more regions;   feeding the one or more regions into a cascade of classifiers, the cascade comprising a plurality of stages; and   removing regions of the image classified as non-text regions from the cascade prior to completion of the cascade to avoid subsequent processing of the removed regions.   
     
     
         2 . The method of  claim 1 , wherein the cascade comprises seven AdaBoost layers. 
     
     
         3 . The method of  claim 2 , wherein each layer of the cascade has an equal or greater number of classifiers than each previous layer of the cascade. 
     
     
         4 . The method of  claim 2 , wherein the classifiers in layers are secondarily ordered based on speed of computation. 
     
     
         5 . The method of  claim 1 , further comprising: outputting an output data comprising identified text regions separated from non-text regions. 
     
     
         6 . The method of  claim 1 , further comprising utilizing a binarization process including:
 classifying individual pixels as one of: non-text, light potential-text, and dark potential-text; and   outputting binarization output data.   
     
     
         7 . The method of  claim 6 , further comprising utilizing two neighborhood thresholds: TLight=μ+kσ and TDark=μ−kσ where and μ and σ are the mean and variance within the selected neighborhood respectively, and k is a constant. 
     
     
         8 . The method of  claim 6 , further comprising:
 grouping the pixels into connected components based on their classification and proximity to other pixels.   
     
     
         9 . The method of  claim 8 , further comprising classifying the connected component as text or non-text based on one or more factors including:
 a number of pixels in the connected component;   a number of pixels on the border of the connected component;   a height of the connected component;   a width of the connected component;   a ratio of the height of the connected component to the width of the connected component;   a ratio of the pixels in the connected component to the width of the connected component multiplied by the height of the connected component; and   a local size of text in the connected component.   
     
     
         10 . The method of  claim 9 , further comprising grouping the connected components into lines of text based on a color distance between colors of two connected components. 
     
     
         11 . The method of  claim 1 , further comprising removing regions classified as text regions from the cascade prior to completion of the cascade when a confidence level exceeds a threshold, wherein the confidence level indicates the likelihood of a region being a text region. 
     
     
         12 . The method of  claim 1 , further comprising:
 receiving training images;   feeding the training images into the cascade;   comparing classifier results to known training image results; and   adapting one or more of an order of stages in the cascade, an order of classifiers in the stages, one or more classifier confidence level thresholds, and the classifiers by selecting features for each classifier that reduce a number of false positive and false negative detections by a reduced number of tests.   
     
     
         13 . A system for detecting text in real-world images comprising:
 a processor including:   a dividing logic to divide an image into one or more regions;   a calculating logic to calculate a cascade of classifiers, the cascade comprising a plurality of stages, each stage including one or more weak classifiers, wherein the plurality of stages is organized to start out with classifiers that are most useful for ruling out non-text regions;   a feeding logic to feed the one or more regions into the cascade remove non-text image regions logic to remove image regions classified as the non-text regions from the cascade prior to completion of the cascade, to avoid subsequent processing of the removed regions.   
     
     
         14 . The system of  claim 13 , further comprising:
 an outputting logic to output an output data comprising identified text regions separated from non-text regions.   
     
     
         15 . The system of  claim 13 , further comprising binarization logic including:
 logic to classify individual pixels as one of: non-text, light potential-text, and dark potential-text.   
     
     
         16 . The system of  claim 13 , further comprising a training system including:
 a feed logic to feed training images into the cascade;   a comparison logic to compare classifier results to known training image results; and   an adapting logic to adapt one or more of an order of stages in cascade of classifiers, an order of classifiers in the stages, one or more classifiers confidence level thresholds, and the classifiers by selecting features for each classifier that reduce the number of false positive and false negative detections by a reduced number of tests.   
     
     
         17 . A training system to detect text in real-world images comprising a processor including:
 a cascade comprising a plurality of stages, each stage including one or more weak classifiers, wherein the plurality of stages is organized to start out with classifiers that are most useful for ruling out non-text regions;   a feed logic to feed training images into the cascade;   a comparison logic to compare classifier results to known training image results; and   an adapting logic to adapt one or more of:
 an order of stages in the cascade of classifiers, 
 an order of classifiers in the stages, 
 one or more classifiers confidence level thresholds, and 
 the classifiers by selecting features for each classifier that reduce the number of false positive and false negative detections by a reduced number of tests.

Join the waitlist — get patent alerts

Track US2011091098A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.