Computer-readable recording medium storing self-supervised training program, method, and device
Abstract
A non-transitory computer-readable recording medium stores a self-supervised training program for causing a computer to execute a process including: generating data that indicates a second molecule obtained by replacing a value that indicates each of a predetermined percentage of atoms among the atoms contained in a first molecule, with zero; acquiring a prediction result by inputting the data that indicates the second molecule to a machine learning model that performs prediction regarding a molecular structure; and updating a parameter of the machine learning model, based on a comparison result between correct answer data that corresponds to the first molecule and the prediction result.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A non-transitory computer-readable recording medium storing a self-supervised training program for causing a computer to execute a process comprising:
generating data that indicates a second molecule obtained by replacing a value that indicates each of a predetermined percentage of atoms among the atoms contained in a first molecule, with zero; acquiring a prediction result by inputting the data that indicates the second molecule to a machine learning model that performs prediction regarding a molecular structure; and updating a parameter of the machine learning model, based on a comparison result between correct answer data that corresponds to the first molecule and the prediction result.
2 . The non-transitory computer-readable recording medium according to claim 1 , wherein the prediction result is information that indicates which of the atoms included in a group of a plurality of predefined types of the atoms at least the atoms replaced from the atoms contained in the first molecule, among the atoms contained in the second molecule, are.
3 . The non-transitory computer-readable recording medium according to claim 1 , wherein the prediction result includes information on all the atoms contained in the second molecule.
4 . The non-transitory computer-readable recording medium according to claim 1 , wherein the machine learning model includes: a first portion that extracts a feature that indicates the molecular structure of the second molecule, from the data that indicates the second molecule; and a second portion that outputs the prediction result according to a specified task, based on the feature extracted by the first portion, and the first portion of the trained machine learning model is used for transfer training.
5 . The non-transitory computer-readable recording medium according to claim 4 , wherein
the second portion includes different output portions that each correspond to an output of the first portion for each of the atoms contained in the second molecule, and the acquiring the prediction result includes acquiring the prediction result for each of the atoms contained in the second molecule at one time.
6 . A self-supervised training method comprising:
generating data that indicates a second molecule obtained by replacing a value that indicates each of a predetermined percentage of atoms among the atoms contained in a first molecule, with zero; acquiring a prediction result by inputting the data that indicates the second molecule to a machine learning model that performs prediction regarding a molecular structure; and updating a parameter of the machine learning model, based on a comparison result between correct answer data that corresponds to the first molecule and the prediction result.
7 . The self-supervised training method according to claim 6 , wherein the prediction result is information that indicates which of the atoms included in a group of a plurality of predefined types of the atoms at least the atoms replaced from the atoms contained in the first molecule, among the atoms contained in the second molecule, are.
8 . The self-supervised training method according to claim 6 , wherein the prediction result includes information on all the atoms contained in the second molecule.
9 . The self-supervised training method according to claim 6 , wherein the machine learning model includes: a first portion that extracts a feature that indicates the molecular structure of the second molecule, from the data that indicates the second molecule; and a second portion that outputs the prediction result according to a specified task, based on the feature extracted by the first portion, and the first portion of the trained machine learning model is used for transfer training.
10 . The self-supervised training method according to claim 9 , wherein
the second portion includes different output portions that each correspond to an output of the first portion for each of the atoms contained in the second molecule, and the acquiring the prediction result includes acquiring the prediction result for each of the atoms contained in the second molecule at one time.
11 . An information processing device comprising:
a memory; and a processor coupled to the memory and configured to: generate data that indicates a second molecule obtained by replacing a value that indicates each of a predetermined percentage of atoms among the atoms contained in a first molecule, with zero; acquire a prediction result by inputting the data that indicates the second molecule to a machine learning model that performs prediction regarding a molecular structure; and update a parameter of the machine learning model, based on a comparison result between correct answer data that corresponds to the first molecule and the prediction result.
12 . The information processing device according to claim 11 , wherein the prediction result is information that indicates which of the atoms included in a group of a plurality of predefined types of the atoms at least the atoms replaced from the atoms contained in the first molecule, among the atoms contained in the second molecule, are.
13 . The information processing device according to claim 11 , wherein the prediction result includes information on all the atoms contained in the second molecule.
14 . The information processing device according to claim 11 , wherein the machine learning model includes: a first portion that extracts a feature that indicates the molecular structure of the second molecule, from the data that indicates the second molecule; and a second portion that outputs the prediction result according to a specified task, based on the feature extracted by the first portion, and the first portion of the trained machine learning model is used for transfer training.
15 . The information processing device according to claim 14 , wherein
the second portion includes different output portions that each correspond to an output of the first portion for each of the atoms contained in the second molecule, and the acquiring the prediction result includes acquiring the prediction result for each of the atoms contained in the second molecule at one time.Join the waitlist — get patent alerts
Track US2024394548A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.