US2006167698A1PendingUtilityA1
System and method for generating an identification signal for electronic devices
Assignee: NELLYMOSER INC A MASSACHUSETTSPriority: Dec 31, 2001Filed: Mar 22, 2006Published: Jul 27, 2006
Est. expiryDec 31, 2021(expired)· nominal 20-yr term from priority
G10H 2240/056G10H 2250/291G10H 2230/021G10H 3/125G10H 2250/285H04M 19/041G10H 2250/235G10H 2250/265
41
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system and method for creating a ring tone for an electronic device takes as input a phrase sung in a human voice and transforms it into a control signal controlling, for example, a ringer on a cellular telephone. Time-varying features of the input signal are analyzed to segment the signal into a set of discrete notes and assigning to each note a chromatic pitch value. The set of note start and stop times and pitches are then translated into a format suitable for controlling the device.
Claims
exact text as granted — not AI-modified1 . A method for generating an identification signal, comprising:
accepting as input a monophonic audio signal of limited duration; translating said monophonic audio signal to a representation of a series of discrete tones; and producing a control signal from said representation of discrete tones, said control signal suitable for causing a transponder to generate a signal, where said generated signal is a translation of said monophonic audio signal; wherein translating said monophonic audio signal to the representation of the series of discrete tones includes segmenting the monophonic audio signal into a series of segments according to time varying features of the audio signal that include a feature associated with energy and a feature associated with spectral composition, wherein each tone in the series of discrete tones is associated with a different segment in the series of segments.
2 . A method for generating an identification signal, comprising:
accepting as input a voice signal of limited duration; translating said voice signal to a representation of a series of discrete tones; and producing a control signal from said representation of discrete tones, said control signal suitable for causing a transponder to generate a signal, where said generated signal is a translation of said voice signal; wherein translating said voice signal to the representation of the series of discrete tones includes segmenting the voice signal into a series of segments according to time varying features of the voice signal that include a feature associated with energy and a feature associated with spectral composition, wherein each tone in the series of discrete tones is associated with a different segment in the series of segments.
3 . The method of claim 2 wherein said generated signal is melodically human-recognizable.
4 . The method of claim 2 wherein said generated signal is rhythmically human-recognizable.
5 . The method of claim 2 wherein accepting as input further comprises receiving said voice signal over a telephone connection.
6 . The method of claim 5 wherein said telephone connection is wireless.
7 . The method of claim 2 wherein said step of accepting as input further comprises receiving said voice signal over a microphone attached to a computer.
8 . The method of claim 2 wherein said translating step further comprises translating said voice signal to a range of tones within the capability of a mobile telephone audio output synthesizer.
9 . The method of claim 2 further comprising the step of transmitting said control signal to a tone-producing output device responsive to said control signal.
10 . The method of claim 2 wherein said translating step further comprises:
generating a digital representation of said voice signal; dividing said digitized signal into a plurality of frames; extracting analysis data from each said frame; and formatting said analysis data into a frame representation.
11 . The method of claim 10 further comprising the step of segmenting said signal by counting instances of increased signal amplitude in said frames, and
for each instance of increased amplitude, determining a change in each of pitch, energy, and spectral composition in a region around said instance of increased amplitude, whereby a segment is defined by a start frame having an instance of increased amplitude and an end frame is defined by changes in pitch, energy and spectral composition in relation to selected thresholds.
12 . The method of claim 10 wherein said translating step further comprises grouping said frames into a plurality of regions.
13 . The method of claim 12 wherein each said region is determined from a count of consecutive upward short-term average change in cepstral-domain energy followed by a count of consecutive downward short-term average change in cepstral-domain energy.
14 . The method of claim 12 further comprising the step of determining the existence of a candidate note start frame in each said region.
15 . The method of claim 13 further comprising the step of determining a candidate note start frame in each said region as the last frame within said region in which the count of consecutive upward short-term average change in cepstral-domain energy is not zero.
16 . The method of claim 14 further comprising the step of determining which regions of said plurality have a valid note start frame.
17 . The method of claim 14 , wherein determining a candidate note start frame further comprises the step of determining if the cepstral domain energy of a particular frame is greater than a cepstral domain energy threshold and a frame immediately before said particular frame was below said cepstral domain energy threshold.
18 . The method of claim 14 , wherein determining a candidate note start frame further comprises the step of determining whether a fundamental frequency range of a particular frame is above a fundamental frequency range threshold and whether an energy range for said particular frame is above an energy range threshold.
19 . The method of claim 14 , further comprising the step of determining a stop frame corresponding to each start frame.
20 . The method of claim 15 , further comprising the step of determining a stop frame by locating the first frame after a start frame in which cepstral energy is below said cepstral domain energy threshold.
21 . The method of claim 20 , further comprising the step of defining the stop frame as a frame between two and ten frames before a subsequent start frame if no frame having cepstral energy below said cepstral domain energy threshold is found.
22 . The method of claim 19 further comprising the step of verifying each start and stop frame pair by determining whether
a) average voicing probability is above a voicing probability threshold, b) average short-time energy is above an average short-time energy threshold, and c) average fundamental frequency is above an average fundamental frequency threshold.
23 . The method of claim 2 wherein the feature associated with energy includes a time-domain energy.
24 . The method of claim 2 wherein the feature associated with energy includes a cepstral-domain energy.
25 . The method of claim 2 wherein the time varying features according to which the voice signal is segmented include at least two features associated with energy.
26 . The method of clam 2 wherein the feature associated with spectral composition includes a cepstral coefficient.
27 . The method of claim 2 wherein the time varying features according to which the voice signal is segmented further include a feature associated with periodicity.
28 . The method of claim 27 wherein the feature associated with periodicity includes a fundamental frequency.
29 . The method of claim 27 wherein feature associated with periodicity includes a voicing probability.
30 . Apparatus for generating an identification signal comprising:
a voice signal receiver; a translator having as its input a voice signal received by said voice signal receiver and having as its output a representation of discrete tones where an audio presentation of said discrete tones would be human-recognizable as a translation of said voice signal; wherein the translator includes an estimation module with outputs of a time varying feature associated with each of energy and spectral composition from the voice signal and a segmentation module responsive to the time varying features with an output of a segmentation of the voice signal into a series of segments according to the time varying features, such that each in the series of output discrete tones is associated with a different segment in the series of segments.
31 . The apparatus of claim 30 wherein said voice signal receiver comprises an analog telephone receiver.
32 . The apparatus of claim 30 wherein said voice signal receiver further comprises a voice-to-digital signal transducer.
33 . The apparatus of claim 30 wherein said voice signal receiver further comprises a recording device.
34 . The apparatus of claim 30 wherein said translator further comprises a feature estimation module to determine values for at least one time-varying feature of said input signal.
35 . The apparatus of claim 34 wherein said translator further comprises a pitch assignment module responsive to signal energy in each segment output by said segmentation module.
36 . The apparatus of claim 34 wherein said feature estimation module further comprises a primary feature module, a secondary feature module and a tertiary feature module.
37 . The apparatus of claim 36 wherein said primary feature module determines a plurality of values for each of time-domain energy, fundamental frequency, cepstral-domain energy, and voicing probability.
38 . The apparatus of claim 35 wherein said segmentation module further comprises a first-phase segmentation module and a second-phase segmentation module.
39 . The apparatus of claim 38 wherein said first-phase segmentation module groups a plurality of successive frames of said input signal into at least one region in response to output of said feature estimation module.
40 . The apparatus of claim 39 wherein said region is a plurality of frames in which a change in energy increases immediately followed by frames in which change in energy decreases.
41 . The apparatus of claim 40 in which said region has a minimum rider of frames.
42 . The apparatus of claim 39 wherein said second-phase segmentation module determines if said at least one region has a valid note start frame and if so, determines a stop frame.
43 . The apparatus of claim 42 wherein said second-phase segmentation module determines said valid note start frame in response to cepstral domain energy by determining whether a frame has a cepstral domain energy greater than a cepstral domain energy threshold preceded by a frame having a cepstral domain energy less than said cepstral domain threshold.
44 . The apparatus of claim 42 wherein said second-phase segmentation module determines a valid note start frame if the fundamental frequency exceeds a fundamental energy threshold and if the non-cepstral domain energy exceeds an energy threshold.
45 . The apparatus of claim 39 further comprising a segmentation post-processor to verify said start and stop frame in response to average voicing probability, average short-time energy, and average fundamental frequency of said start and stop frame.
46 . The apparatus of claim 35 wherein said pitch assignment module assigns an integer between 32 and 83, said integer corresponding to the MIDI note number for pitch.
47 . The apparatus of claim 35 wherein said pitch assignment module comprises an intranote pitch assignment subsystem and an internote pitch assignment subsystem.
48 . The apparatus of claim 47 wherein said internote pitch assignment subsystem corrects pitches determined by said intranote pitch assignment subsystem.
49 . The apparatus of claim 48 wherein said internote pitch assignment subsystem further comprises a key finding stage to assign a scale to a note sequence output by said intranote pitch assignment subsystem.
50 . The apparatus of claim 48 wherein said internote pitch assignment subsystem further comprises a pairwise correction stage to examine a pitch and its preceding pitch for conformity to voice-leading rules,
if a pair is determined to be dissonant according to said voice-leading rules, the internote pitch assignment subsystem corrects the pitches of said pair if the pitch adjustment does not cause dissonance in an adjacent pair.Join the waitlist — get patent alerts
Track US2006167698A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.