Systems and methods for self-guided sequence selection and extrapolation
Abstract
Embodiments described herein provide systems and methods for training a sequential recommendation model. Methods include determining a difficulty and quality (DQ) score associated with user behavior sequences from a training dataset. User behavior sequences are sampled during training based on their DQ scores. A meta-extrapolator may also be trained based on user behavior sequences sampled according to DQ score. The meta-extrapolator may be trained with high quality low difficulty sequences. The meta-extrapolator may then be used with an input of high quality high difficulty sequences to generate synthetic user behavior sequences. The synthetic user behavior sequences may be used to augment the training dataset to fine-tune the sequential recommendation model, while continuing to sample user behavior sequences based on DQ score. As the DQ score is based on current model predictions, DQ scores iteratively update during the training process.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for training a sequential recommendation model, comprising:
receiving, via a communication interface, a training dataset comprising a plurality of user behavior sequences; determining, for at least one user behavior sequence from the training dataset, a difficulty score and a quality score based on a first prediction generated by the sequential recommendation model from the at least one user behavior sequence; determining a respective combined difficulty and quality (DQ) score for each of the plurality of user behavior sequences based on the respective difficulty and quality scores; sampling one or more user behavior sequences from the plurality of user behavior sequences based on the respective DQ scores; generating, by the sequential recommendation model, a second prediction from the sampled one or more user behavior sequences; and updating parameters of the sequential recommendation model based on a training objective comparing the second prediction and a ground truth corresponding to the sampled one or more user behavior sequences.
2 . The method of claim 1 , wherein the difficulty score of the at least one user behavior sequence is based on an accuracy of next-item predictions of the sequential recommendation model for the at least one user behavior sequence.
3 . The method of claim 1 , wherein the difficulty score of the at least one user behavior sequence is combined with a previous difficulty score of the at least one user behavior sequence.
4 . The method of claim 1 , wherein the quality score of the at least one user behavior sequence is based on a measure of variance of prediction scores of the sequential recommendation model across items in the at least one user behavior sequence.
5 . The method of claim 1 , wherein the quality score of the at least one user behavior sequence is combined with a previous quality score of the at least one user behavior sequence.
6 . The method of claim 1 , wherein the determining the respective combined DQ score of the at least one user behavior sequence comprises summing a square of the difficulty score with a square of the quality score.
7 . The method of claim 6 , wherein the determining the respective combined DQ score of the at least one user behavior sequence further comprises raising the sum of the squares to a predetermined power.
8 . The method of claim 1 , wherein the determining the respective combined DQ score the at least one user behavior sequence comprises weighting the difficulty score and the quality score differently.
9 . The method of claim 1 , wherein the difficulty score and the quality score are iteratively updated based on the updated sequential recommendation model.
10 . A system for training a sequential recommendation model, comprising:
a memory that stores the sequential recommendation model; a communication interface that receives a plurality of user behavior sequences; and one or more hardware processors that:
receives, via the communication interface, a training dataset comprising a plurality of user behavior sequences;
determines, for at least one user behavior sequence from the training dataset, a difficulty score and a quality score based on predictions of the sequential recommendation model;
determines a respective combined difficulty and quality (DQ) score for each of the plurality of user behavior sequences based on the respective difficulty and quality scores;
samples one or more user behavior sequences from the plurality of user behavior sequences based on the respective DQ scores;
updates parameters of the sequential recommendation model based on a training objective comparing predictions generated by the sequential recommendation model and ground-truth from the sampled one or more user behavior sequences; and fine-tunes the sequential recommendation model based at least in part on the sampled one or more user behavior sequences.
11 . The system of claim 10 , wherein the difficulty score of the at least one user behavior sequence is based on an accuracy of next-item predictions of the sequential recommendation model for the at least one user behavior sequence.
12 . The system of claim 10 , wherein the difficulty score of the at least one user behavior sequence is combined with a previous difficulty score of the at least one user behavior sequence.
13 . The system of claim 10 , wherein the quality score of the at least one user behavior sequence is based on a measure of variance of prediction scores of the sequential recommendation model across items in the at least one user behavior sequence.
14 . The system of claim 10 , wherein the quality score of the at least one user behavior sequence is combined with a previous quality score of the at least one user behavior sequence.
15 . The system of claim 10 , wherein the determining the respective combined DQ score of the at least one user behavior sequence comprises summing a square of the difficulty score with a square of the quality score.
16 . The system of claim 15 , wherein the determining the respective combined DQ score of the at least one user behavior sequence further comprises raising the sum of the squares to a predetermined power.
17 . The system of claim 10 , wherein the determining the respective combined DQ score the at least one user behavior sequence comprises weighting the difficulty score and the quality score differently.
18 . A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:
receiving, via a communication interface, a training dataset comprising a plurality of user behavior sequences; determining a first set of difficulty and quality (DQ) scores, corresponding to the plurality of user behavior sequences, respectively, based on predictions generated by a sequential recommendation model during training; updating parameters of the sequential recommendation model based on a training objective comparing predictions generated by the sequential recommendation model and ground-truth from the plurality of user behavior sequences sampled based on respective DQ scores; selecting a first subset of user behavior sequences and a second subset of user behavior sequences from the plurality of user behavior sequences based on the first set of DQ scores; training an extrapolator using the first subset of user behavior sequences; generating, by the trained extrapolator, a plurality of synthetic user behavior sequences based on the second subset of user behavior sequences; and fine-tuning the sequential recommendation model using the plurality of synthetic user behavior sequences.
19 . The non-transitory machine-readable medium of claim 18 , wherein the first subset of user behavior sequences includes user behavior sequences with a respective quality score above a first predetermined threshold, and a respective difficulty score below a second predetermined threshold.
20 . The non-transitory machine-readable medium of claim 18 , wherein the second subset of user behavior sequences includes user behavior sequences with a respective quality score above a first predetermined threshold, and a respective difficulty score above a second predetermined threshold.Join the waitlist — get patent alerts
Track US2024070744A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.