User behavior data analysis method and device
Abstract
A user behavior data analysis method and device, used to accurately analyze user behavior and make advertising more targeted. The method comprises: obtaining behavior data generated in a data source after a user is registered with the data source ( 101 ), the data source containing behavior data respectively generated by all users registered with the data source, and the behavior data being data information recording the behavior of a user in the data source; extracting a user label from the behavior data of the user generated in the data source ( 102 ), the user label being information indicative of user behavior; obtaining preset directed population characteristics ( 103 ), the directed population characteristics being characteristics possessed by the population meeting the directed characteristics requirement; according to the behavior data of the user generated in the data source and the user label, extracting a target user group complying with the directed population characteristics from all users in the data source ( 104 ), the target user group comprising a plurality of users complying with the directed population characteristics.
Claims
exact text as granted — not AI-modified1 . A method for analyzing user behavior data, comprising:
obtaining behavior data generated by a user in a data source after the user registers with the data source, wherein the data source comprises behavior data generated by each user that registers with the data source and the behavior data is data information recording a behavior of a user in the data source; extracting a user tag from the behavior data generated by the user in the data source, wherein the user tag is information representing a behavior of the user; obtaining a preset oriented audience characteristic, wherein the oriented audience characteristic is a characteristic of an audience meeting an oriented characteristic requirement; and extracting a target user group meeting the oriented audience characteristic from all users in the data source, based on the behavior data generated by the user in the data source and the user tag, wherein the target user group comprises multiple users meeting the oriented audience characteristic, wherein extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the user in the data source and the user tag comprises: extracting an oriented category from classified categories in the data source based on the oriented audience characteristic; performing statistics to determine the number of user behaviors, each of which with the user tag meeting the oriented category, in the data source; and extracting users, each of which with the number of the user behaviors exceeding an oriented category threshold, in the data source, to form the target user group, wherein the target user group comprises all users each of which with the number of the user behaviors exceeding the oriented category threshold.
2 . (canceled)
3 . The method according to claim 1 , wherein performing statistics to determine the number of the user behaviors, each of which with the user tag meeting the oriented category, in the data source comprises:
calculating the number of the user behaviors, each of which with the user tag meeting the oriented category, in the data source by using the following formula:
number=Σ i=1 N (λ i *Σ j=1 M count j );
wherein number is the number of the user behaviors, N is the number of data sources, λ i is a weight of an i-th data source, M is the number of oriented categories in the i-th data source, and count j is the number of user behaviors of a user in a j-th oriented category in each data source.
4 . The method according to claim 1 , wherein extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the user in the data source and the user tag comprises:
obtaining a keyword of the oriented audience characteristic based on the oriented audience characteristic; matching the keyword with the extracted user tag, and calculating the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source; calculating an oriented audience score of each user having a user behavior with the user tag being matched with the keyword successfully in the data source, based on a forgetting factor and the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source; and extracting users, each of which with the oriented audience score exceeding an oriented audience correlation threshold, in the data source, to form the target user group, wherein the target user group comprises all users, each of which with the oriented audience score exceeding the oriented audience correlation threshold, in the data source.
5 . The method according to claim 4 , wherein after obtaining the keyword of the oriented audience characteristic based on the oriented audience characteristic, the method further comprises:
obtaining a filter word which is related to the keyword but is not matched with the oriented audience characteristic, based on the obtained keyword; and wherein matching the keyword with the extracted user tag and calculating the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source, comprises: matching the keyword and the filter word with the extracted user tag respectively; and calculating the number of all user behaviors, each of which with the user tag being matched with the keyword successfully but failing to be matched with the filter word, in the data source.
6 . The method according to claim 4 , wherein calculating the oriented audience score of each user having a user behavior with the user tag being matched with the keyword successfully in the data source based on the forgetting factor and the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source, comprises:
calculating the oriented audience score of each user having a user behavior with the user tag being matched with the keyword successfully in the data source, by using the following formula:
score
=
1
1
+
γ
*
exp
[
-
∑
begin_time
end_time
∑
i
=
1
N
(
λ
i
*
S
i
*
F
(
x
)
)
/
b
]
;
wherein score is the oriented audience score, N is the number of data sources, λ i is a weight of an i-th data source, S i is the number of user behaviors, each of which with the user tag being matched with the keyword successfully, in the i-th data source, F(X) is the forgetting factor,
F
(
X
)
=
-
log
2
(
cur
-
est
)
hl
,
cur is a current time when calculating score, est is a time when the user behavior is generated, hl is a half-life period, begin_time is a start time of the behavior data recorded in the data source, end_time is an end time of the behavior data recorded in the data source, γ is a control parameter for a range of the oriented audience score, and b is a control parameter for an increment speed of the oriented audience score.
7 . The method according to claim 1 , wherein extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the user in the data source and the user tag comprises:
selecting a training sample set from all users in the data source based on the oriented audience characteristic; extracting a behavior characteristic from a user tag of a user in the training sample set, wherein a characteristic value of the behavior characteristic is a term frequency-inverse document frequency (TF-IDF) of a word representing the behavior characteristic; training a categorization model with the behavior characteristic using a categorization method; and categorizing all users in the data source by the categorization model, to obtain the target user group, wherein the target user group comprises all users screened out by the categorization model.
8 . The method according to claim 7 , wherein the TF-IDF is calculated by using the following formula:
TFIDF
=
tf
(
t
,
d
)
*
log
2
(
N
n
i
+
0.01
)
∑
[
tf
(
t
,
d
)
*
log
2
(
N
n
i
+
0.01
)
]
2
,
wherein tf(t,d) is the number of user behaviors in the data source, t is a word representing the behavior characteristic, d is the behavior data in the data source, N is the number of user behaviors of all users, and n i is the number of user behaviors of a user selected as the training sample set.
9 . The method according to claim 1 , wherein after extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the user in the data source and the user tag, the method further comprises:
obtaining an audience characteristic distribution of all users in the target user group; and filtering out a user in the target user group exceeding a characteristic distribution range of the audience characteristic distribution, to obtain a first corrected target user group, wherein the first corrected target user group comprises users in the target user group within the characteristic distribution range of the audience characteristic distribution.
10 . The method according to claim 1 , wherein after extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the use in the data source and the user tag, the method further comprises:
updating the behavior data generated by the user in the data source; and correcting the target user group meeting the oriented audience characteristic based on the updated behavior data, to obtain a second corrected target user group.
11 . The method according to claim 10 , wherein correcting the target user group meeting the oriented audience characteristic based on the updated behavior data to obtain the second corrected target user group comprises:
extracting an updated user tag from the updated behavior data, and extracting multiple users meeting the oriented audience characteristic based on the updated behavior data and the updated user tag, to form the second corrected target user group.
12 . The method according to claim 1 , wherein after extracting the target user group meeting the oriented audience characteristic from all users in the data source based on the behavior data generated by the user in the data source and the user tag, the method further comprises:
verifying a correlation between multiple users in the target user group and the oriented audience characteristic; correcting behavior data in a data source corresponding to a user, of which the correlation is less than a correlation threshold, in the target user group; and correcting the target user group meeting the oriented audience characteristic based on the corrected behavior data, to obtain a third corrected target user group.
13 . The method according to claim 12 , wherein correcting the target user group meeting the oriented audience characteristic based on the corrected behavior data to obtain the third corrected target user group comprises:
extracting a corrected user tag from the corrected behavior data, and extracting multiple users meeting the oriented audience characteristic based on the corrected behavior data and the corrected user tag, to form the third corrected target user group.
14 . A device for analyzing user behavior data, comprising:
a data obtaining processor, configured to obtain behavior data generated by a user in a data source after the user registers with the data source, wherein the data source comprises behavior data generated by each user that registers with the data source and the behavior data is data information recording a behavior of a user in the data source; a tag extraction processor, configured to extract a user tag from the behavior data generated by the user in the data source, wherein the user tag is information representing a behavior of the user; a characteristic obtaining processor, configured to obtain a preset oriented audience characteristic, wherein the oriented audience characteristic is a characteristic of an audience meeting an oriented characteristic requirement; and a user group extraction processor, configured to extract a target user group meeting the oriented audience characteristic from all users in the data source, based on the behavior data generated by the user in the data source and the user tag, wherein the target user group comprises multiple users meeting the oriented audience characteristic, wherein the user group extraction processor comprises: an oriented category extraction sub-processor, configured to extract an oriented category from classified categories in the data source based on the oriented audience characteristic; a first user behavior statistic sub-processor, configured to perform statistics to determine the number of user behaviors, each of which with the user tag meeting the oriented category, in the data source; and a first user group extraction sub-processor, configured to extract users, each of which with the number of the user behaviors exceeding an oriented category threshold, in the data source, to form the target user group, wherein the target user group comprises all users each of which with the number of the user behaviors exceeding the oriented category threshold.
15 . (canceled)
16 . The device according to claim 14 , wherein the first user behavior statistic sub-processor is configured to calculate the number of the user behaviors, each of which with the user tag meeting the oriented category, in the data source by using the following formula:
number=Σ i=1 N (λ i *Σ j=1 M count j );
wherein number is the number of the user behaviors, N is the number of data sources, λ i is a weight of an i-th data source, M is the number of oriented categories in the i-th data source, and count j is the number of user behaviors of a user in a j-th oriented category in each data source.
17 . The device according to claim 14 , wherein the user group extraction processor comprises:
a keyword obtaining sub-processor, configured to obtain a keyword of the oriented audience characteristic based on the oriented audience characteristic; a second user behavior statistic sub-processor, configured to match the keyword with the extracted user tag, and calculate the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source; an audience score calculation sub-processor, configured to calculate an oriented audience score of each user having a user behavior with the user tag being matched with the keyword successfully in the data source, based on a forgetting factor and the number of all user behaviors, each of which with the user tag being matched with the keyword successfully, in the data source; and a second user group extraction sub-processor, configured to extract users, each of which with the oriented audience score exceeding an oriented audience correlation threshold, in the data source, to form the target user group, wherein the target user group comprises all users, each of which with the oriented audience score exceeding the oriented audience correlation threshold, in the data source.
18 . The device according to claim 17 , wherein the user group extraction processor further comprises a filter word obtaining sub-processor, wherein
the filter word obtaining sub-processor is configured to obtain a filter word which is related to the keyword but is not matched with the oriented audience characteristic, based on the obtained keyword; and the second user behavior statistic sub-processor is configured to match the keyword and the filter word with the extracted user tag respectively; and calculate the number of all user behaviors, each of which with the user tag being matched with the keyword successfully but failing to be matched with the filter word, in the data source.
19 . The device according to claim 17 , wherein the audience score calculation sub-processor is configured to calculate the oriented audience score of each user having a user behavior with the user tag being matched with the keyword successfully in the data source, by using the following formula:
score
=
1
1
+
γ
*
exp
[
-
∑
begin_time
end_time
∑
i
=
1
N
(
λ
i
*
S
i
*
F
(
x
)
)
/
b
]
;
wherein score is the oriented audience score, N is the number of data sources, λ i is a weight of an i-th data source, S i is the number of user behaviors, each of which with the user tag being matched with the keyword successfully, in the i-th data source, F(X) is the forgetting factor,
F
(
X
)
=
-
log
2
(
cur
-
est
)
hl
,
cur is a current time when calculating score, est is a time when the user behavior is generated, hl is a half-life period, begin_time is a start time of the behavior data recorded in the data source, end_time is an end time of the behavior data recorded in the data source, γ is a control parameter for a range of the oriented audience score, and b is a control parameter for an increment speed of the oriented audience score.
20 . The device according to claim 19 , wherein the user group extraction processor comprises:
a sample selection sub-processor, configured to select a training sample set from all users in the data source based on the oriented audience characteristic; a behavior characteristic extraction sub-processor, configured to extract a behavior characteristic from a user tag of a user in the training sample set, wherein a characteristic value of the behavior characteristic is a term frequency-inverse document frequency (TF-IDF) of a word representing the behavior characteristic; a model train sub-processor, configured to a categorization model with the behavior characteristic using a categorization method; and a user categorization sub-processor, configured to categorize all users in the data source by the categorization model, to obtain the target user group, wherein the target user group comprises all users screened out by the categorization model.
21 . The device according to claim 20 , wherein the TF-IDF of the behavior characteristic extracted by the behavior characteristic extraction sub-processor is calculated by using the following formula:
TFIDF
=
tf
(
t
,
d
)
*
log
2
(
N
n
i
+
0.01
)
∑
[
tf
(
t
,
d
)
*
log
2
(
N
n
i
+
0.01
)
]
2
,
wherein tf(t,d) is the number of user behaviors in the data source, t is a word representing the behavior characteristic, d is the behavior data in the data source, N is the number of user behaviors of all users, and n i is the number of user behaviors of a user selected as the training sample set.
22 . The device according to claim 14 , wherein the device for analyzing user behavior data further comprises:
a characteristic distribution obtaining processor, configured to obtain an audience characteristic distribution of all users in the target user group; and a first user group correction processor, configured to filter out a user in the target user group exceeding a characteristic distribution range of the audience characteristic distribution, to obtain a first corrected target user group, wherein the first corrected target user group comprises users in the target user group within the characteristic distribution range of the audience characteristic distribution.
23 . The device according to claim 14 , wherein the device for analyzing user behavior data further comprises:
a behavior data update processor, configured to update the behavior data generated by the user in the data source; and a second user group correction processor, configured to correct the target user group meeting the oriented audience characteristic based on the updated behavior data, to obtain a second corrected target user group.
24 . The device according to claim 23 , wherein the second user group correction processor is configured to extract an updated user tag from the updated behavior data, and extracting multiple users meeting the oriented audience characteristic based on the updated behavior data and the updated user tag, to form the second corrected target user group.
25 . The device according to claim 14 , wherein the device for analyzing user behavior data further comprises:
a correlation verification processor, configured to verify a correlation between multiple users in the target user group and the oriented audience characteristic; a behavior data correction processor, configured to correct behavior data in a data source corresponding to a user, of which the correlation is less than a correlation threshold, in the target user group; and a third user group correction processor, configured to correct the target user group meeting the oriented audience characteristic based on the corrected behavior data, to obtain a third corrected target user group.
26 . The device according to claim 25 , wherein the third user group correction processor is configured to extract a corrected user tag from the corrected behavior data, and extract multiple users meeting the oriented audience characteristic based on the corrected behavior data and the corrected user tag, to form the third corrected target user group.Join the waitlist — get patent alerts
Track US2016379268A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.