Method and device for obtaining video clip, server, and storage medium
Abstract
The present application belongs to the technical field of audio and video, and relates to a method and device for obtaining a video clip, a server, and a storage medium. The method includes in response to obtaining a clip in live stream video data of a performance live stream room, using audio data from the live stream video data and audio data of an original performer to determine a target timepoint parameter of the live stream video data. The method includes obtaining a target video clip according to a start timepoint and an end timepoint in the target timepoint parameter. The present application is used to capture a more complete video clip.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for obtaining a video clip, comprising:
obtaining live streaming video data of a performance live streaming room; determining target time point pairs of the live streaming video data based on audio data of the live streaming video data and audio data of an original performer, wherein each of the target time point pairs comprises a start time point and an end time point; obtaining a candidate target video clip from the live streaming video data based on the target time point pairs; and determining the candidate target video clip as a target video clip in response to at least one of:
an amount of gift resources of the candidate target video clip exceeding a gift resource threshold, or
in response to an amount of comment information of the candidate target video clip exceeding a comment information threshold, or
in response to an amount of like information of the candidate target video clip exceeding a like information threshold.
2 . The method according to claim 1 , wherein said determining the target time point pairs of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer comprises:
determining first time points of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer; and determining the target time point pairs corresponding to the first time points centered at the first time points based on a preset interception time duration.
3 . The method according to claim 2 , wherein the audio data of the live streaming video data is audio data of a song sung by a host, and the audio data of the original performer is audio data of the song sung by the original singer;
said determining the first time points of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer comprises: obtaining lyrics of the song by performing voice recognition on the audio data of the live streaming video data; obtaining the audio data of the song sung by the original singer based on the lyrics; determining a lyric similarity between audio features of the audio data of the song sung by the original singer and audio features of the audio data of the live streaming video data for each sentence of the lyrics; and determining a time point corresponding to a position in the lyrics with a highest lyric similarity above a lyric information threshold, as the first time point of the live streaming video data.
4 . The method according to claim 2 , further comprising:
determining second time points of the live streaming video data based on interaction information of accounts other than a host account of the live streaming video data; wherein said determining the target time point pairs corresponding to the first time points centered at the first time points based on the preset interception time duration comprises: determining a time point of the first time points as a target time point in response to determination that the time point belongs to the second time points, and deleting the time point in response to determination that the time point does not belong to the second time points; and determining the target time point pair corresponding to the target time point centered at the target time point based on the preset interception time duration.
5 . The method according to claim 4 , wherein said determining the second time points in the live streaming video data based on interaction information of accounts other than a host account of the live streaming video data comprises:
determining a middle time point or an end time point of a first time period as the second time point of the live streaming video data in response to an amount of gift resources of the live streaming video data in the first time period exceeding another gift resource threshold; determining a middle time point or an end time point of a second time period as the second time point of the live streaming video data in response to an amount of comment information of the live streaming video data in the second time period exceeding another comment information threshold; or, determining a middle time point or an end time point of a third time period as the second time point of the live streaming video data in response to a number of likes of the live streaming video data in the third time period exceeding another like information threshold.
6 . The method according to claim 5 , further comprising:
obtaining a number of each type of recognized gift images by recognizing gift images in the live streaming video data for the first time period; and determining the amount of the gift resources in the first time period based on the number of each type of gift images.
7 . The method according to claim 1 , further comprising:
replacing a first end time point with a second end time point and deleting a second start time point and the second end time point in the target time point pairs in response to a first start time point being earlier than the second start time point, the first end time point being earlier than the second end time point, and the second start time point being earlier than the first end time point, wherein the first end time point corresponds to the first start time point, the second end time point corresponds to the second start time point, the first start time point and the second start time point are different and are comprised in the target time point pairs.
8 . The method according to claim 1 , further comprising:
generating link information of the target video clip; and sending the link information to login terminals of other accounts than a host account in the performance live streaming room for displaying the link information on a playback interface or a live streaming end interface of the performance live streaming room of the login terminals of other accounts.
9 . A device for obtaining a video clip, comprising:
a processor; and a memory for storing instructions executable by the processor; wherein, the processor is configured to perform operations comprising: obtaining live streaming video data of a performance live streaming room; determining target time point pairs of the live streaming video data based on audio data of the live streaming video data and audio data of an original performer, wherein each of the target time point pairs comprises a start time point and an end time point; obtaining a candidate target video clip from the live streaming video data based on the target time point pairs; and determining the candidate target video clip as a target video clip in response to an amount of gift resources of the candidate target video clip exceeding a gift resource threshold, or in response to an amount of comment information of the candidate target video clip exceeding a comment information threshold, or in response to an amount of like information of the candidate target video clip exceeding a like information threshold.
10 . The device according to claim 9 , wherein said determining the target time point pairs of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer comprises:
determining first time points of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer; and determining the target time point pairs corresponding to the first time points centered at the first time points based on a preset interception time duration.
11 . The device according to claim 10 , wherein the audio data of the live streaming video data is audio data of a song sung by a host, and the audio data of the original performer is audio data of the song sung by the original singer;
said determining the first time points of the live streaming video data based on the audio data of the live streaming video data and the audio data of the original performer comprises: obtaining lyrics of the song by performing voice recognition on the audio data of the live streaming video data; obtaining the audio data of the song sung by the original singer based on the lyrics; determining a lyric similarity between audio features of the audio data of the song sung by the original singer and audio features of the audio data of the live streaming video data for each sentence of the lyrics; and determining a time point data corresponding to a position in the lyrics with a highest lyric similarity above a lyric information threshold, as the first time point of the live streaming video data.
12 . The device according to claim 10 , wherein the operations further comprise:
determining second time points of the live streaming video data based on interaction information of accounts other than a host account of the live streaming video data; wherein said determining the target time point pairs corresponding to the first time points centered at the first time points based on the preset interception time duration comprises: determining a time point of the first time points as a target time point in response to determination that the time point belongs to the second time points, and deleting the time point in response to determination that the time point does not belong to the second time points; and determining the target time point pair corresponding to the target time point centered at the target time point based on the preset interception time duration.
13 . The device according to claim 12 , wherein said determining the second time points in the live streaming video data based on interaction information of accounts other than a host account of the live streaming video data comprises:
determining a middle time point or an end time point of a first time period as the second time point of the live streaming video data in response to an amount of gift resources of the live streaming video data in the first time period exceeding another gift resource threshold; determining a middle time point or an end time point of a second time period as the second time point of the live streaming video data in response to an amount of comment information of the live streaming video data in the second time period exceeding another comment information third threshold; or, determining a middle time point or an end time point of a third time period as the second time point of the live streaming video data in response to a number of likes of the live streaming video data in the third time period exceeding another like information threshold.
14 . The device according to claim 13 , wherein the operations further comprise:
obtaining a number of each type of recognized gift images by recognizing gift images in the live streaming video data for the first time period; and determining the amount of the gift resources in the first time period based on the number of each type of gift images.
15 . The device according to claim 9 , wherein the operations further comprise:
replacing a first end time point with a second end time point and deleting a second start time point and the second end time point in the target time point pairs in response to a first start time point being earlier than the second start time point, the first end time point being earlier than the second end time point, and the second start time point being earlier than the first end time point, wherein the first end time point corresponds to the first start time point, the second end time point corresponds to the second start time point, the first start time point and the second start time point are different and are comprised in the target time point pairs.
16 . The device according to claim 9 , wherein the operations further comprise:
generating link information of the target video clip; and sending the link information to login terminals of other accounts than a host account in the performance live streaming room for displaying the link information on a playback interface or a live streaming end interface of the performance live streaming room of the login terminals of other accounts.
17 . A non-transitory computer-readable storage medium having stored thereon instructions which, when being executed by a processor of a server, cause the server to perform operations comprising:
obtaining live streaming video data of a performance live streaming room; determining target time point pairs of the live streaming video data based on audio data of the live streaming video data and audio data of an original performer, wherein each of the target time point pairs comprises a start time point and an end time point; obtaining a candidate target video clip from the live streaming video data based on the target time point pairs; and determining the candidate target video clip as a target video clip in response to an amount of gift resources of the candidate target video clip exceeding a gift resource threshold, or in response to an amount of comment information of the candidate target video clip exceeding a comment information threshold, or in response to an amount of like information of the candidate target video clip exceeding a like information threshold.Join the waitlist — get patent alerts
Track US2022303644A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.