Efficient extraction of intelligence from web data
Abstract
Embodiments are directed to a system for gathering and processing web data. The system provides an expression-based social media monitoring (SMM) tool that pulls from a world wide web an initial data universe that includes web data relevant to a targeted index that has been identified by an entity as being of importance to said entity. An initial set of themes relevant to the targeted index is pulled from the initial data universe, and an expression-based, cognitive data analysis tool codes the initial data universe under the initial set of relevant themes to filter portions of the initial data universe that fall under the initial set of relevant themes and portions of the initial data universe that do not fall under the initial set of relevant themes.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of gathering and processing web data, the method comprising:
manually identifying a targeted index based on a top-level inquiry identified by an entity as being of importance to said entity; pulling from a world wide web an initial data universe comprising web data relevant to said targeted index; discovering from said initial data universe an initial set of themes relevant to said targeted index; coding said initial data universe under said initial set of relevant themes to filter portions of said initial data universe that fall under said initial set of relevant themes and portions of said initial data universe that do not fall under said initial set of relevant themes; analyzing said portions of said initial data universe that do not fall under said initial set of relevant themes to identify any additional themes relevant to said targeted index; and in the event that said additional relevant themes are identified, further coding under said additional relevant themes said portions of said initial data universe that do not fall under said initial set of relevant themes; wherein said targeted index is defined by said coded portions of said initial data universe that fall under said initial set of relevant themes, along with said coded portions of said data universe that fall under said additional relevant themes.
2 . The method of claim 1 wherein said further coding under said additional relevant themes is repeated until an accuracy confidence level for said additional relevant themes meets or exceeds a threshold.
3 . The method of claim 1 wherein said threshold comprises 90 percent confidence level with a +/−5% standard error.
4 . The method of claim 1 further comprising defining said targeted index by other data relevant to said targeted index but not pulled from the world wide web.
5 . The method of claim 1 wherein said other data comprises enterprise data about said entity.
6 . The method of claim 1 wherein said pulling, said discovering, said coding, said analyzing and said further coding are periodically repeated to capture changes in said initial data universe.
7 . The method of claim 1 further comprising conveying insights about said targeted index said top-level inquiry identified by said entity as being of importance to said entity by representing said insights visually on a user interface.
8 . The method of claim 1 wherein said insights comprise discovered stories derived from:
said coded portions of said initial data universe that fall under said initial set of relevant themes; and
said coded portions of said initial data universe that fall under said additional relevant themes.
9 . The method of claim 1 wherein:
a keyword-based SMM search tool performs said pulling from the world wide web an initial data universe; and
an automated spreadsheet performs said coding by applying automated coding to said initial data universe under said initial set of relevant themes to filter portions of said initial data universe that fall under said initial set of relevant themes and portions of said initial data universe that do not fall under said initial set of relevant themes.
10 . The method of claim 9 wherein:
an expression-based SMM search tool performs said pulling from the world wide web an initial data universe; and
an expression-based, cognitive data analysis tool performs said applying automated coding to said initial data universe under said initial set of relevant themes to filter portions of said initial data universe that fall under said initial set of relevant themes and portions of said initial data universe that do not fall under said initial set of relevant themes.Join the waitlist — get patent alerts
Track US2016070805A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.