High precision set expansion for large concepts
Abstract
A set expansion system is described herein that improves precision, recall, and performance of prior set expansion methods for large sets of data. The system maintains high precision and recall by 1) identifying the quality of particular lists and applying that quality through a weight, 2) allowing for the specification or negative examples in a set of seeds to reduce the introduction of bad entities into the set, and 3) applying a cutoff to eliminate lists that include a low number of positive matches. The system may perform multiple passes to first generate a good candidate result set and then refine the set to find a set with highest quality. The system may also apply Map Reduce or other distributed processing techniques to allow calculation in parallel. Thus, the system efficiently expands large concept sets from a potentially small set of initial seeds from readily available web data.
Claims
exact text as granted — not AI-modifiedI/We claim:
1 . A computer-implemented method to measure the quality of a candidate result set expanded from a set of seed items, the method comprising:
receiving one or more seed items that represent members of a concept set for which a user wants to automatically generate additional members; receiving one or more lists that include some items that are members of the concept set and other items that are not members of the concept set; receiving a candidate result set that expands the received seed items to include items suspected of being members of the concept set; determining a weight for each received list based on the received seeds, wherein the weight corresponds to an initial measure of the quality of the list; determining a similarity metric of each item in the received candidate result set with the received seed items based on which of the received lists contain each item and the determined list weights; determining a quality of the received candidate result set by combining the determined similarity metrics; and outputting the determined quality of the candidate result set, wherein the preceding steps are performed by at least one processor.Join the waitlist — get patent alerts
Track US2017124206A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.