Token stream differencing with moved-block detection
Abstract
Methods and apparatus implementing systems and techniques for differencing token streams and detecting moved blocks of tokens. In general, in one implementation, the technique includes: obtaining a first token stream and a second token stream, comparing the first and second token streams to identify a group of tokens that are substantially similar in the first and second token streams, the similar-tokens group including common sub-sequences, which are identical in the first and second token streams, and at least one unmatched token, and presenting matched token information corresponding to the similar-tokens group to represent changes in document flow.
Claims
exact text as granted — not AI-modified1 - 41 . (canceled)
42 . A method comprising:
identifying sequences of text units that are shared by first text information and second text information; grouping the shared sequences of text units into one or more matched blocks of text based on a length of one or more intervening text units, which separate the shared sequences of text units in one of the first text information or the second text information, relative to a length of the shared sequences of text units; and generating a representation of the one or more matched blocks of text to show a result of the grouping.
43 . The method of claim 42 , wherein the grouping comprises grouping the shared sequences of text units into multiple matched text blocks, each including one or more intervening text units, and grouping the matched text blocks based on an order of the matched text blocks in the first text information and the second text information.
44 . The method of claim 43 , wherein grouping the matched text blocks comprises grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks.
45 . The method of claim 44 , wherein grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks comprises selecting the matched text blocks for grouping based on lengths of the shared sequences of text units.
46 . The method of claim 43 , wherein generating the representation of the one or more matched blocks of text comprises presenting matched block information to a user.
47 . The method of claim 42 , wherein the text units comprise words in a language, and the first text information and the second text information correspond to a reading order for the words.
48 . The method of claim 42 , wherein grouping the shared sequences of text units into one or more matched blocks of text based on a length comprises comparing lengths with respect to a difference threshold and a minimum length threshold.
49 . The method of claim 48 , wherein the difference threshold comprises a configurable difference threshold configured based on a nature of the text units.
50 . A machine-readable medium tangibly storing a software product comprising instructions operable to cause one or more programmable processors to perform operations comprising:
identifying sequences of text units that are shared by first text information and second text information; grouping the shared sequences of text units into one or more matched blocks of text based on a length of one or more intervening text units, which separate the shared sequences of text units in one of the first text information or the second text information, relative to a length of the shared sequences of text units; and generating a representation of the one or more matched blocks of text to show a result of the grouping.
51 . The machine-readable medium of claim 50 , wherein the grouping comprises grouping the shared sequences of text units into multiple matched text blocks, each including one or more intervening text units, and grouping the matched text blocks based on an order of the matched text blocks in the first text information and the second text information.
52 . The machine-readable medium of claim 51 , wherein grouping the matched text blocks comprises grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks.
53 . The machine-readable medium of claim 52 , wherein grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks comprises selecting the matched text blocks for grouping based on lengths of the shared sequences of text units.
54 . The machine-readable medium of claim 51 , wherein generating the representation of the one or more matched blocks of text comprises presenting matched block information to a user.
55 . The machine-readable medium of claim 50 , wherein the text units comprise words in a language, and the first text information and the second text information correspond to a reading order for the words.
56 . The machine-readable medium of claim 50 , wherein grouping the shared sequences of text units into one or more matched blocks of text based on a length comprises comparing lengths with respect to a difference threshold and a minimum length threshold.
57 . The machine-readable medium of claim 56 , wherein the difference threshold comprises a configurable difference threshold configured based on a nature of the text units.
58 . A system comprising:
one or more processors; and a machine-readable storage device operable to cause the one or more processors to perform operations comprising: identifying sequences of text units that are shared by first text information and second text information, grouping the shared sequences of text units into one or more matched blocks of text based on a length of one or more intervening text units, which separate the shared sequences of text units in one of the first text information or the second text information, relative to a length of the shared sequences of text units, and generating a representation of the one or more matched blocks of text to show a result of the grouping.
59 . The system of claim 58 , wherein the grouping comprises grouping the shared sequences of text units into multiple matched text blocks, each including one or more intervening text units, and grouping the matched text blocks based on an order of the matched text blocks in the first text information and the second text information.
60 . The system of claim 59 , wherein grouping the matched text blocks comprises grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks.
61 . The system of claim 60 , wherein grouping the matched text blocks and one or more of the shared sequences of text units that are not included in any of the matched text blocks comprises selecting the matched text blocks for grouping based on lengths of the shared sequences of text units.
62 . The system of claim 59 , wherein generating the representation of the one or more matched blocks of text comprises presenting matched block information to a user.
63 . The system of claim 58 , wherein the text units comprise words in a language, and the first text information and the second text information correspond to a reading order for the words.
64 . The system of claim 58 , wherein grouping the shared sequences of text units into one or more matched blocks of text based on a length comprises comparing lengths with respect to a difference threshold and a minimum length threshold.
65 . The system of claim 64 , wherein the difference threshold comprises a configurable difference threshold configured based on a nature of the text units.Join the waitlist — get patent alerts
Track US2009012777A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.