System and methods for bid optimization in real-time bidding
Abstract
A method of operating a demand side platform (DSP) includes determining a current state of the DSP, wherein the current state of the DSP is based on a remaining bid budget and remaining number of opportunities, receiving, at the DSP, a bid request for one or more advertisement impressions, and determining an uncertainty of a predicted user response probability. The method further includes determining a risk tendency value based on the current state of the DSP, determining an adjusted value of the one or more advertisement impressions based on the uncertainty and risk tendency, determining a bid price for each of the one or more advertisement impressions based on the adjusted value of the one or more advertisement impressions, transmitting the bid price to an exchange platform to participate in an auction, receiving an auction result and updating the current state of the DSP based on the auction result.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of operating a demand side platform (DSP), the method comprising:
determining a current state of the DSP, wherein the current state of the DSP is based on a remaining bid budget and remaining number of opportunities; receiving, at the DSP, a bid request for one or more advertisement impressions; determining an uncertainty of a predicted user response probability; determining a risk tendency value based on the current state of the DSP; determining an adjusted value of the one or more advertisement impressions based on the uncertainty and risk tendency; determining a bid price for each of the one or more advertisement impressions based on the adjusted value of the one or more advertisement impressions; transmitting the bid price to an exchange platform to participate in an auction; receiving an auction result; and updating the current state of the DSP based on the auction result.
2 . The method of claim 1 , wherein the bid price is further determined based on a reinforcement learning trained model.
3 . The method of claim 1 , wherein determining the risk tendency value comprises:
determining a sign of the risk tendency value; determining a monotonicity of the risk tendency value; and determining applicability of an early state approximation.
4 . The method of claim 1 , wherein determining the risk tendency value comprises:
training a multi-layer perceptron to learn a risk tendency function associating the risk tendency value with current values of remaining bid budget and remaining number of opportunities.
5 . The method of claim 4 , wherein training the multi-layer perceptron comprises adding Gaussian noise to the risk tendency function during training.
6 . The method of claim 4 , wherein training the multi-layer perceptron comprises populating and updating an experience buffer comprising a set of DSP state data associated with leading values of a reward function.
7 . The method of claim 1 , further comprising:
receiving by the DSP, from an external device, via a network, at least one of a configuration command enabling prediction uncertainty compensation or a configuration command enabling one or more risk tendency compensation modes.
8 . A demand side platform (DSP), the DSP comprising:
a processor; a network interface; and a memory containing instructions, which when executed by the processor, cause the DSP to:
determine a current state of the DSP, wherein the current state of the DSP is based on a remaining bid budget and remaining number of opportunities,
receive, via the network interface, a bid request for one or more advertisement impressions,
determine an uncertainty of a predicted user response probability,
determine a risk tendency value based on the current state of the DSP,
determine an adjusted value of the one or more advertisement impressions based on the uncertainty and risk tendency,
determine a bid price for each of the one or more advertisement impressions based on the adjusted value of the one or more advertisement impressions,
transmit, via the network interface, the bid price to an exchange platform to participate in an auction,
receive, via the network interface, an auction result, and
update the current state of the DSP based on the auction result.
9 . The DSP of claim 8 , wherein the bid price is further determined based on a reinforcement learning trained model.
10 . The DSP of claim 8 , wherein determining the risk tendency value comprises:
determining a sign of the risk tendency value; determining a monotonicity of the risk tendency value; and determining applicability of an early state approximation.
11 . The DSP of claim 8 , wherein determining the risk tendency value comprises:
training a multi-layer perceptron to learn a risk tendency function associating the risk tendency value with current values of remaining bid budget and remaining number of opportunities.
12 . The DSP of claim 11 , wherein training the multi-layer perceptron comprises adding Gaussian noise to the risk tendency function during training.
13 . The DSP of claim 11 , wherein training the multi-layer perceptron comprises populating and updating an experience buffer comprising a set of DSP state data associated with leading values of a reward function.
14 . The DSP of claim 8 , wherein the memory further contains instructions, which, when executed by the processor, cause the DSP to:
receive by the DSP, from an external device, via the network interface, at least one of a configuration command enabling prediction uncertainty compensation or a configuration command enabling one or more risk tendency compensation modes.
15 . A non-transitory, computer-readable medium containing instructions, which when executed by a processor, cause a demand side platform (DSP) to:
determine a current state of the DSP, wherein the current state of the DSP is based on a remaining bid budget and remaining number of opportunities, receive, via a network interface, a bid request for one or more advertisement impressions, determine an uncertainty of a predicted user response probability, determine a risk tendency value based on the current state of the DSP, determine an adjusted value of the one or more advertisement impressions based on the uncertainty and risk tendency, determine a bid price for each of the one or more advertisement impressions based on the adjusted value of the one or more advertisement impressions, transmit, via the network interface, the bid price to an exchange platform to participate in an auction, receive, via the network interface, an auction result, and update the current state of the DSP based on the auction result.
16 . The non-transitory, computer-readable medium of claim 15 , wherein the bid price is further determined based on a reinforcement learning trained model.
17 . The non-transitory, computer-readable medium of claim 15 , wherein determining the risk tendency value comprises:
determining a sign of the risk tendency value; determining a monotonicity of the risk tendency value; and determining applicability of an early state approximation.
18 . The non-transitory, computer-readable medium of claim 15 , wherein determining the risk tendency value comprises:
training a multi-layer perceptron to learn a risk tendency function associating the risk tendency value with current values of remaining bid budget and remaining number of opportunities.
19 . The non-transitory, computer-readable medium of claim 18 , wherein training the multi-layer perceptron comprises adding Gaussian noise to the risk tendency function during training.
20 . The non-transitory, computer-readable medium of claim 18 , wherein training the multi-layer perceptron comprises populating and updating an experience buffer comprising a set of DSP state data associated with leading values of a reward function.Join the waitlist — get patent alerts
Track US2023089895A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.