US2020387398A1PendingUtilityA1
Fuzzy matching for computing resources
Est. expiryMay 3, 2037(~10.8 yrs left)· nominal 20-yr term from priority
Inventors:Yu Xia
G06F 9/30018H04L 41/085G06F 16/3334G06F 9/5027G06F 16/2246H04L 41/0859G06F 16/3344G06F 9/30021G06F 16/90344G06F 9/4881
58
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems and methods are disclosed for normalizing strings to identify computing resources and improve performance and utilization of computing resources. For example, methods may include determining a comparison length based on lengths of strings in a set of strings; padding a first string from the set of strings to the comparison length to obtain a padded string; receiving a second string; determining a distance between the second string and the padded string; and identifying a match between the first string and the second string based on the distance.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A system operable to normalize a plurality of strings associated with respective computing resources of a computing network, the system comprising:
a memory; and a processor, wherein the memory includes instructions executable by the processor to cause the system to:
determine a comparison length for the plurality of strings based on a length of each of the plurality of strings;
in response to the plurality of strings having respective lengths less than the comparison length, pad the plurality of strings using one or more randomly generated characters to increase the respective lengths to the comparison length;
receive a first string associated with an unidentified computing resource;
determine respective counts of character operations to be performed to transform characters of the first string to characters of each of the plurality of strings;
identify a match between one of the plurality of strings and the first string based on the respective count of character operations performed; and
determine the first string to be associated with a known computing resource associated with the one of the plurality strings instead of the unidentified computing resource based on the match.
2 . The system of claim 1 , wherein the memory includes instructions executable by the processor to cause the system to:
responsive to the match, modify the first string based on the one of the plurality of strings from the plurality of strings to obtain a normalized string; and store, display, or transmit the normalized string.
3 . The system of claim 1 , wherein the instructions to determine the comparison length include instructions executable by the processor to cause the system to:
determine the comparison length as a weighted average of strings of the plurality of strings.
4 . The system of claim 1 , wherein the instructions to determine the comparison length include instructions executable by the processor to cause the system to:
rank software publishers associated with strings of the plurality of strings based on indications of an extent or frequency of installations of software published by the software publishers; determine respective weights for strings of the plurality of strings based on a rank of a software publisher associated with the respective string; and determine a weighted average of lengths of strings of the plurality of strings using the respective weights for the strings of the plurality of strings.
5 . The system of claim 1 , wherein the instructions to pad each of the plurality of strings to the comparison length include instructions executable by the processor to cause the system to:
append one or more of the one or more randomly generated characters to the plurality of strings.
6 . The system of claim 1 , wherein the instructions to pad the first string to the comparison length include instructions executable by the processor to cause the system to:
append one or more escaped characters to the first string, wherein the one or more escaped characters do not to occur in the plurality of strings.
7 . The system of claim 1 , wherein the instructions to determine the respective counts include instructions executable by the processor to cause the system to:
determine the respective counts using a dynamic programming algorithm to determine a Damerau-Levenshtein distance between the first string and each of the plurality of strings.
8 . The system of claim 1 , wherein the instructions to identify the match include instructions executable by the processor to cause the system to:
compare the respective counts to a threshold; and identify the match between the one of the plurality of strings and the first string responsive to the respective count being below the threshold.
9 . The system of claim 1 , wherein the memory includes instructions executable by the processor to cause the system to:
update the plurality of strings by adding one or more strings to or removing one or more strings from the plurality of strings; and responsive to the update, determine the comparison length based on lengths of strings in the plurality of strings.
10 . The system of claim 1 , wherein each of the plurality of strings are associated with a record for an installed software, the first string is associated with a software usage record, and the memory includes instructions executable by the processor to cause the system to:
responsive to the match, uninstall the installed software.
11 . A method for normalizing a plurality of strings associated with respective computing resources of a computing network, the method comprising:
determining a comparison length for the plurality of strings based on a length of each of the plurality of strings; in response to the plurality of strings having respective lengths less than the comparison length, padding the plurality of strings using one or more randomly generated characters to increase the respective lengths to the comparison length; receiving a first string associated with an unidentified computing resource; determining respective counts of character operations to be performed to transform characters of the first string to characters of each of the plurality of strings; identifying a match between one of the plurality of strings and the first string based on the respective count of character operations performed; and determining the first string to be associated with a known computing resource associated with the one of the plurality of strings instead of the unidentified computing resource based on the match.
12 . The method of claim 11 , comprising:
responsive to the match, modifying the first string based on the one of the plurality of strings to obtain a normalized string; and storing, displaying, or transmitting the normalized string.
13 . The method of claim 11 , wherein determining the comparison length comprises:
determining the comparison length as a weighted average of respective weights associated with each of the plurality of strings, wherein the respective weights are based on an installation frequency of a computing resource of the respective computing resources described by a string.
14 . The method of claim 11 , wherein determining the comparison length comprises:
ranking software publishers associated with strings of the plurality of strings based on indications of an extent or frequency of installations of software published by the software publishers; determining respective weights for strings of the plurality of strings based on the rank of a software publisher associated with the respective string; and determining a weighted average of lengths of strings of the plurality of strings using the respective weights for the strings of the plurality of strings.
15 . The method of claim 11 , wherein padding each of the plurality of strings to the comparison length comprises:
appending one or more of the randomly generated characters to the plurality of strings.
16 . The method of claim 11 , wherein padding the first string to the comparison length comprises:
appending one or more escaped characters to the plurality of strings, wherein the one or more escaped characters do not to occur in the plurality of strings.
17 . The method of claim 11 , wherein determining the respective counts comprises:
using a dynamic programming algorithm to determine a Damerau-Levenshtein distance between the first string and each of the plurality of strings.
18 . The method of claim 11 , comprising:
responsive to the match, uninstalling an installed software, wherein each of the plurality of strings are associated with a record for the installed software and the first string is associated with a software usage record.
19 . A system operable to facilitate matching of a plurality of strings associated with respective computing resources of a computing network, the system comprising:
a memory; and a processor, wherein the memory includes instructions executable by the processor to cause the system to:
determine a comparison length for the plurality of strings based on a length of each of the plurality of strings;
in response to the plurality of strings having respective lengths less than the comparison length, pad the plurality of strings using one or more randomly generated characters to increase the respective lengths to a comparison length;
receive a first string associated with an unidentified computing resource;
determine respective counts of character operations to be performed to transform characters of the first string to characters of each of the plurality of strings;
identify a match between one of the plurality of strings and the first string based on the respective count of character operations performed; and
determine the first string to be associated with a known computing resource of the respective computing resources associated with the one of the plurality of strings instead of the unidentified computing resource based on the match.
20 . The system of claim 19 , wherein each of the plurality of strings are associated with a record for an installed software, the first string is associated with a software usage record, and the memory includes instructions executable by the processor to cause the system to:
responsive to the match, uninstall the installed software.Join the waitlist — get patent alerts
Track US2020387398A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.