US2015172299A1PendingUtilityA1
Indexing and retrieval of blogs
Est. expirySep 13, 2025(expired)· nominal 20-yr term from priority
G06F 17/30386H04L 63/14G06F 16/951G06F 16/313G06F 16/24
43
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system may receive a feed associated with a blog. The system may extract information from the feed and the blog and create a hybrid document based on the extracted information. The system may further use the hybrid document to determine a relevance of the blog to a search query.
Claims
exact text as granted — not AI-modified1 - 28 . (canceled)
29 . A method performed by one or more server devices, the method comprising:
extracting, by one or more processors associated with the one or more server devices, first information from a post to a blog that includes a plurality of posts; extracting, from the blog and by one or more processors associated with the one or more server devices, second information, associated with the blog, from a source different than the posts to the blog, where the second information is not included in the posts to the blog; and comparing, by one or more processors associated with the one or more server devices, one or more terms of a search query to one or more terms included in the first information and the second information to determine a relevance of the post to the search query.
30 . The method of claim 29 , where the source includes a Rich Site Summary (RSS) feed, an Atom feed, or another structured representation of information contained in the blog.
31 . The method of claim 29 , where the extracting the second information includes:
extracting at least one of a title of the blog, a name of an author of the blog, or a profile of the author of the blog.
32 . The method of claim 29 , where the source includes one or more documents to which the blog links.
33 . The method of claim 29 , where the extracting the first information includes identifying at least one of:
a timestamp that indicates when the post was created, a timestamp that indicates when the post was updated, content of the post, a title of the post, or a name of an author of the post.
34 . The method of claim 29 , further comprising:
comparing the first information to the second information, and determining, based on the comparing, whether at least one of the post or the blog are spam.
35 . The method of claim 34 , where the determining includes determining that the at least one of the post or the blog are spam when at least a portion of the first information does not match a corresponding portion of the second information.
36 . A computer-readable memory device storing computer-executable instructions, the computer-executable instructions comprising:
one or more instructions to extract first information from a particular post to a blog that includes a plurality of posts; one or more instructions to extract, from the blog, second information, associated with the blog, from a source different than the posts to the blog, where the second information is not included in the posts to the blog; and one or more instructions to index another post to the blog based on a content of the other post, the first information, and the second information.
37 . The computer-readable memory device of claim 36 , where the one or more instructions to extract the second information include:
one or more instructions to extract the second information from one or more documents to which the blog links.
38 . The computer-readable memory device of claim 36 , where the second information includes at least one of:
information identifying a geographical location of an author of the blog, information identifying an age of the author of the blog, or information identifying a gender of the author of the blog.
39 . The computer-readable memory device of claim 36 , where the computer-executable instructions further comprise:
one or more instructions to use the index to determine a relevance of the other post to a received search query.
40 . The computer-readable memory device of claim 36 , where the source includes a Rich Site Summary (RSS) feed, an Atom feed, or another structured representation of information contained in the blog.
41 . The computer-readable memory device of claim 36 , where the second information includes at least one of:
information identifying a title of the blog, information identifying a name of an author of the blog, or information identifying a profile of the author of the blog.
42 . The computer-readable memory device of claim 36 , where the first information includes at least one of:
a timestamp that indicates when the post was created, a timestamp that indicates when the post was updated, content of the post, a title of the post, or a name of an author of the post.
43 . The computer-readable memory device of claim 36 , further comprising:
one or more instructions to compare the first information to the second information, and one or more instructions to determine, based on the comparing, whether at least one of the particular post, the other post, or the blog are spam.
44 . A method performed by one or more server devices, the method comprising:
identifying, by one or more processors associated with the one or more server devices, a blog that is associated with:
one or more posts, and
data that is from a source other than the one or more posts;
fetching, by one or more processors associated with the one or more server devices, first information from a particular post of the one or more posts; extracting, by one or more processors associated with the one or more server devices, second information from the data that is from a source other than the one or more posts; determining that the first information is different from the second information, and classifying the particular post as spam in response to the determining; and reducing a rating of the blog post in response to classifying the particular post as spam.
45 . The method of claim 44 , where the determining includes determining that the first information is different from the second information when at least a portion of the first information does not match a corresponding portion of the second information.
46 . The method of claim 44 , where the source includes a Rich Site Summary (RSS) feed, an Atom feed, or another structured representation of information contained in the blog.
47 . The method of claim 44 , where the source includes one or more documents to which the blog links.
48 . A system, comprising:
a memory to store computer-executable instructions; and one or more processors to execute the computer-executable instructions to:
extract first information from a post to a blog that includes a plurality of posts;
extract second information, associated with the blog, from a source that is within the blog and that is different than the posts to the blog, where the second information is not included in the posts to the blog; and
compare one or more terms of a search query to one or more terms included in the first information and the second information to determine a relevance of the post to the search query.Join the waitlist — get patent alerts
Track US2015172299A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.