Automated data modeling for abbreviations utilizing fuzzy reasoning logic
Abstract
A method includes analyzing an enterprise data warehouse to determine name attribute scores based on occurrences of enterprise terms and abbreviations in the enterprise data warehouse. The method also includes generating a scoring summary of phrases, applying fuzzy reasoning logic to identify one or more relationship patterns and weights for the phrases including at least one shared word to produce training data for a data model associated with the enterprise data warehouse, and updating the data model with a new abbreviated field name associated with a new field name based on identifying a closest match of the new field name with the training data.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method, comprising:
analyzing an enterprise data warehouse to determine a plurality of name attribute scores based on a number of occurrences of a plurality of enterprise terms and abbreviations in the enterprise data warehouse using an enterprise abbreviation list that maps the enterprise terms to the abbreviations; generating a scoring summary of a plurality of phrases comprising at least one shared word based on the enterprise terms identified in the enterprise data warehouse and the name attribute scores; applying fuzzy reasoning logic to identify one or more relationship patterns and weights for the phrases comprising at least one shared word based on the scoring summary to produce a plurality of training data for a data model associated with the enterprise data warehouse; and updating the data model with a new abbreviated field name associated with a new field name based on identifying a closest match of the new field name with the training data.
2 . The method of claim 1 , wherein the enterprise data warehouse comprises a plurality of databases and metadata defining one or more aspects of data within the databases.
3 . The method of claim 2 , wherein the databases comprise at least one difference in field formats.
4 . The method of claim 1 , further comprising:
receiving a request to update the data model to include the new field name; and searching the enterprise data warehouse to confirm whether an existing field name matches the new field name.
5 . The method of claim 4 , further comprising:
determining whether a field attribute change is requested based on matching the new field name with the existing field name; and updating a data definition language file associated with the data model based on identifying the field attribute change.
6 . The method of claim 1 , further comprising:
receiving a data source identifier; performing a database schema comparison to identify one or more differences between a source database and a target database based on determining that the data source identifier indicates a database source type; identifying a field name and a data type associated with the one or more differences; checking the field name associated with the one or more differences with respect to a plurality of formatting rules; applying an abbreviation to the field name based on the training data and the formatting rules; and generating a data definition language file for the target database based on the field name, the abbreviation of the field name, and the data type.
7 . The method of claim 1 , further comprising:
receiving a data source identifier; parsing a mapping document to identify one or more differences with respect to an existing mapping document based on determining that the data source identifier indicates a mapping document source type; identifying a field name and a data type associated with the one or more differences; determining whether an existing abbreviation of the field name is found in the enterprise abbreviation list; creating an abbreviation of the field name based on identifying a matching entry in the enterprise abbreviation list; creating the abbreviation of the field name based on the training data in response to a failure to identify the matching entry in the enterprise abbreviation list; and generating a data definition language file for a target database based on the field name, the abbreviation of the field name, and the data type.
8 . The method of claim 7 , further comprising:
identifying a plurality of word groupings in the field name; forming a plurality of combinations of the word groupings for abbreviating the field name; and determining the abbreviation of the field name based on matching at least one of the combinations with the enterprise abbreviation list.
9 . The method of claim 8 , further comprising:
determining the scoring summary for the combinations of the word groupings; and selecting the abbreviation of the field name based on the scoring summary for the combinations of the word groupings.
10 . The method of claim 7 , further comprising:
checking the field name associated with the one or more differences with respect to a plurality of formatting rules; and applying the abbreviation to the field name based on the training data and the formatting rules.
11 . The method of claim 1 , further comprising:
generating a mapping document comprising a source to target mapping of at least one field name, description, and data type for an update of the data model at a target database of the enterprise data warehouse; and deriving a field name based on the description in response to determining that the field name is missing in the target database.
12 . The method of claim 11 , further comprising:
identifying a plurality of enterprise domains associated with different instances of the enterprise abbreviation list; and determining an abbreviation of the field name based on comparing abbreviation data across the enterprise domains.
13 . The method of claim 1 , further comprising:
deriving a plurality of abbreviation options by applying a plurality of abbreviation rules to a phrase based on a failure to locate a corresponding abbreviation in the enterprise abbreviation list; and selecting one of the abbreviation options as a derived abbreviation based on confirming that the derived abbreviation is unique with respect to the enterprise abbreviation list.
14 . The method of claim 13 , further comprising:
comparing the derived abbreviation to a maximum character length limit; parsing the derived abbreviation into a plurality of abbreviated words based on determining that the derived abbreviation exceeds the maximum character length limit; analyzing the enterprise data warehouse to determine a plurality of derived abbreviation scores based on a number of occurrences of the abbreviated words and combinations of the abbreviated words in the enterprise data warehouse; and modifying the derived abbreviation to drop one or more characters based on the derived abbreviation scores.
15 . The method of claim 13 , further comprising:
outputting the abbreviation options to a user interface; confirming the derived abbreviation based on a selection through the user interface; and updating the enterprise abbreviation list with the derived abbreviation based on a request received through the user interface.
16 . The method of claim 1 , further comprising:
receiving a source database selection and a source schema selection associated with the data model through a user interface; receiving a target database selection and a target schema selection associated with the data model through the user interface; and generating a data definition language file based on one or more updates to the data model to map a change between from the source database selection and the source schema selection to the target database selection and the target schema selection.Join the waitlist — get patent alerts
Track US2023169375A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.