Network data mining to determine user interest
Abstract
Mining information from network data traffic to determine interests of online network users is provided herein. A data packet received at a network interface device can be accessed and inspected at line rate speeds. Source or addressing information in the data packet can be extracted to identify an initiating and/or receiving device. The packet can be inspected to identify occurrences of keywords or data features related with one or more subject matters. A vector can be defined for a network device that indicates a relative rank of interest in various subject matters. Furthermore, statistical analysis can be implemented on data stored in one or more interest vectors to determine information pertinent to network user interests. The information can facilitate providing value-added products or services to network users.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
obtaining source or destination information from a data packet; comparing data within the data packet to one or more keywords, the keywords are associated with a common subject of interest; establishing a number of times in which one of the keywords matches at least a portion of the data; ranking an interest in the common subject in part based on the established number of matches; and defining a user interest vector that correlates the source or destination information, the common subject and the ranked interest.
2 . The method of claim 1 , further comprising comparing the common subject with respect to an additional subject of interest based in part on the ranking.
3 . The method of claim 1 , further comprising:
accessing an additional data packet that comprises the identified source or addressing information; identifying instances in which one of the keywords matches a portion of data within the additional data packet; and updating the ranked interest based in part on a number of times that one of the keywords is found within the additional data packet.
4 . The method of claim 3 , further comprising:
recording a time when each match to data within the data packet is established; recording a time when each match to data within the additional data packet is identified; and determining a change in the ranked interest as a function of at least one or more recorded times.
5 . The method of claim 1 , further comprising employing machine learning to decompose a user interest vector that contains a ranked interest for each of a plurality of subject matters.
6 . The method of claim 5 , further comprising identifying one or more of the plurality of subject matters that have a predetermined probability of being associated with a single user or device.
7 . The method of claim 1 , further comprising:
aggregating ranked interests associated with multiple user interest vectors; and identifying a group of device users that share interest in the common subject based in part on the aggregation.
8 . A switch, comprising:
an analysis component that obtains source or destination information from a received data packet; an inspection component that updates an occurrence value each time that an interest indicator is identified within the data packet; the interest indicator is correlated with a subject; and an interest compilation component that defines a user interest vector having a user identity field and a user interest field, the user identity field includes the obtained source or destination information and the user interest field couples the subject with the updated occurrence value.
9 . The switch of claim 8 , further comprising an interest categorization component that ranks the subject based at least on the updated occurrence value.
10 . The switch of claim 8 , further comprising an aggregation component that clusters the user interest vector with at least one additional user interest vector based in part on the updated occurrence value.
11 . The switch of claim 8 , further comprising a reference component that compiles a list of synonyms pertinent to the interest indicator, the inspection component updates the occurrence value each time the interest indicator matches an entry in the list of synonyms.
12 . The switch of claim 8 , further comprising a time stamp component that records an update time for each instance that the inspection component updates the occurrence value.
13 . The switch of claim 12 , further comprising an interest monitoring component that determines a frequency with which the occurrence value is updated and ascertains a degree of interest based in part on the determined update frequency.
14 . The switch of claim 13 , further comprising an interest evolution component that analyzes changes in the determined update frequency, the interest monitoring component employs the analyzed changes in part to ascertain the degree of interest.
15 . The switch of claim 8 , further comprising a query engine that receives a request for data, inspects the user interest vector for the requested data and provides a response to the request.
16 . The switch of claim 15 , wherein the requested data includes at least one of:
a number of users having at least a threshold interest in the subject; a number of users having at least the threshold interest in the subject during a period of time; source or destination information of a cluster of user interest vectors, wherein each of the cluster of user interest vectors comprises a ranked occurrence value pertinent to the subject; or a time of day in which the occurrence value is updated substantially at a threshold update frequency.
17 . The switch of claim 8 , the inspection component compares the interest indicator to data within the data packet by employing substantially line rate deep packet inspection.
18 . The switch of claim 8 , wherein:
the interest compilation component defines the user interest vector to have a user interest field for each of a plurality of subjects of interest; the inspection component compares the data packet to at least one interest indicator correlated with each of the plurality of subjects of interest; and the inspection component updates an interest counter associated with a particular subject of interest when an interest indicator correlated with the particular subject of interest matches data within the data packet.
19 . The switch of claim 18 , further comprising a user parsing component that employs machine learning to distinguish a subject of interest from the plurality of subjects of interest that is attributable to an individual user.
20 . The switch of claim 18 , further comprising an interest parsing component that analyzes changes in the user interest vector over a threshold period and identifies a dominant user based on a dominant or persistent subject of interest.
21 . The switch of claim 18 , further comprising an artificial intelligence component that decomposes data in the user interest vector and identifies a number of potential users associated with the data.
22 . The switch of claim 8 , the interest compilation component generates at least one additional user interest vector during an established period of time, or is associated with one or more selected network devices.
23 . A system, comprising:
means for accessing a data packet; means for identifying source or addressing information within the data packet; means for identifying instances where one or more of a plurality of keywords match data within the data packet, wherein the plurality of keywords relate to a common subject; means for ranking an interest in the common subject based in part on the identified number of instances; and means for defining a user interest vector that correlates the source or addressing information, the common subject and the ranked interest.Join the waitlist — get patent alerts
Track US2013318015A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.