Deprecated: The each() function is deprecated. This message will be suppressed on further calls in /home/zhenxiangba/zhenxiangba.com/public_html/phproxy-improved-master/index.php on line 456
JP2018146901A - Acoustic analyzing method and acoustic analyzing apparatus - Google Patents
[go: Go Back, main page]

JP2018146901A - Acoustic analyzing method and acoustic analyzing apparatus - Google Patents

Acoustic analyzing method and acoustic analyzing apparatus Download PDF

Info

Publication number
JP2018146901A
JP2018146901A JP2017044432A JP2017044432A JP2018146901A JP 2018146901 A JP2018146901 A JP 2018146901A JP 2017044432 A JP2017044432 A JP 2017044432A JP 2017044432 A JP2017044432 A JP 2017044432A JP 2018146901 A JP2018146901 A JP 2018146901A
Authority
JP
Japan
Prior art keywords
sound
signal
sound piece
signals
acoustic
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP2017044432A
Other languages
Japanese (ja)
Other versions
JP6841095B2 (en
Inventor
陽 前澤
Akira Maezawa
陽 前澤
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Yamaha Corp
Original Assignee
Yamaha Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Yamaha Corp filed Critical Yamaha Corp
Priority to JP2017044432A priority Critical patent/JP6841095B2/en
Publication of JP2018146901A publication Critical patent/JP2018146901A/en
Application granted granted Critical
Publication of JP6841095B2 publication Critical patent/JP6841095B2/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Images

Landscapes

  • Circuit For Audible Band Transducer (AREA)

Abstract

PROBLEM TO BE SOLVED: To estimate a time sequence of a plurality of sound pieces configuring an acoustic signal, based on the acoustic signal.SOLUTION: An acoustic analyzing apparatus 100 estimates a time sequence of a plurality of sound piece signals Q configuring an acoustic signal S, by comparing each of N sound piece signals Q with the acoustic signal S. A correlation analyzing unit 22 calculates a cross-correlations of the sound piece signal Q and the acoustic signal S, in each of N sound piece signals Q. An estimation processing unit 24 estimates a time sequence of the plurality of sound piece signals Q configuring the acoustic signal S. The estimation processing unit 24 estimates the time sequence of the plurality of sound piece signals Q, so that an accumulated cross correlation for two or more sound piece signals Q is maximized, in a cross correlation of the sound piece signal Q and the acoustic signal S at an end of the two or more sound piece signal Q arranged on a time axis.SELECTED DRAWING: Figure 1

Description

本発明は、音を表す音響信号を解析する技術に関する。   The present invention relates to a technique for analyzing an acoustic signal representing sound.

事前に用意された複数の音片を時間軸上で相互に配列することにより多様な音響信号を合成する音響処理技術(例えば録音編集方式の音声合成技術)が従来から提案されている。例えば特許文献1には、事前に録音された音片と規則合成処理で生成された音声とを相互に結合することで、合成音声を生成する技術が開示されている。   Conventionally, an acoustic processing technique (for example, a voice synthesis technique of a recording and editing method) that synthesizes various acoustic signals by arranging a plurality of pieces of sound prepared in advance on the time axis has been proposed. For example, Patent Document 1 discloses a technique for generating synthesized speech by mutually combining a previously recorded sound piece and a speech generated by a rule synthesis process.

特開2006−145691号公報JP 2006-145691 A

音響信号の合成に適用された複数の音片の時系列を、合成後の音響信号から推定することが要求される場面がある。例えば、電車等の交通機関の案内音声を合成する場面では、合成後の音響信号を構成する複数の音片を推定することで、案内音声の発話内容に応じた各種の情報を利用者に提供するサービスが実現される。以上の事情を考慮して、本発明は、音響信号を構成する複数の音片の時系列を当該音響信号から推定することを目的とする。   There is a scene where it is required to estimate the time series of a plurality of sound pieces applied to the synthesis of an acoustic signal from the synthesized acoustic signal. For example, in a scene where guidance voices of transportation such as trains are synthesized, various information corresponding to the utterance contents of the guidance voice is provided to the user by estimating a plurality of sound pieces constituting the synthesized acoustic signal. Service is realized. In view of the above circumstances, an object of the present invention is to estimate a time series of a plurality of sound pieces constituting an acoustic signal from the acoustic signal.

以上の課題を解決するために、本発明の好適な態様に係る音響解析方法は、N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する。また、本発明の好適な態様に係る音響解析装置は、N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する解析処理部を具備する。   In order to solve the above problems, an acoustic analysis method according to a preferred aspect of the present invention provides a plurality of sound pieces constituting the sound signal by comparing each of the N sound piece signals with the sound signal. Estimate the time series of the signal. The acoustic analysis device according to a preferred aspect of the present invention estimates a time series of a plurality of sound piece signals constituting the sound signal by comparing each of the N sound piece signals with the sound signal. An analysis processing unit is provided.

本発明の第1実施形態における音響解析装置の構成図である。It is a lineblock diagram of an acoustic analysis device in a 1st embodiment of the present invention. 音響信号と複数の音片信号との関係を示す説明図である。It is explanatory drawing which shows the relationship between an acoustic signal and a several sound piece signal. 音響信号と音片信号との相互相関の説明図である。It is explanatory drawing of the cross correlation of an acoustic signal and a sound piece signal. 推定処理部の動作の説明図である。It is explanatory drawing of operation | movement of an estimation process part. 音響解析装置の動作を例示するフローチャートである。It is a flowchart which illustrates operation | movement of an acoustic analyzer. 第3実施形態における情報提供装置の構成図である。It is a block diagram of the information provision apparatus in 3rd Embodiment.

<第1実施形態>
図1は、本発明の第1実施形態に係る音響解析装置100の構成図である。音響解析装置100は、音声を表す音響信号Sを解析する信号処理装置であり、制御装置12と記憶装置14とを具備するコンピュータシステムで実現される。例えば携帯電話機、スマートフォンまたはパーソナルコンピュータ等の各種の情報処理装置が音響解析装置100として利用され得る。
<First Embodiment>
FIG. 1 is a configuration diagram of an acoustic analysis apparatus 100 according to the first embodiment of the present invention. The acoustic analysis device 100 is a signal processing device that analyzes an acoustic signal S representing voice, and is realized by a computer system including a control device 12 and a storage device 14. For example, various information processing apparatuses such as a mobile phone, a smartphone, or a personal computer can be used as the acoustic analysis apparatus 100.

制御装置12は、例えばCPU(Central Processing Unit)等の処理回路で構成され、音響解析装置100の動作を統括的に制御する。記憶装置14は、制御装置12が実行するプログラムと制御装置12が使用する各種のデータとを記憶する。例えば磁気記録媒体および半導体記録媒体等の公知の記録媒体が記憶装置14として利用され得る。相互に別体で構成された同種または異種の複数の記録媒体の組合せを記憶装置14として利用することも可能である。   The control device 12 is configured by a processing circuit such as a CPU (Central Processing Unit), for example, and comprehensively controls the operation of the acoustic analysis device 100. The storage device 14 stores a program executed by the control device 12 and various data used by the control device 12. For example, a known recording medium such as a magnetic recording medium and a semiconductor recording medium can be used as the storage device 14. A combination of a plurality of recording media of the same type or different types configured separately from each other can be used as the storage device 14.

第1実施形態の記憶装置14は、音響信号Sを記憶する。図2に例示される通り、音響信号Sは、複数の音片信号Q(Q〜Q)を時系列に配列して相互に接続することで事前に生成された時間領域の信号である。音片信号Qは、言語音を構成する部分的な音声(以下「音片」という)の波形を表す信号である。1個の音片は、例えば単語,文節,語句等の分節単位を発音した音声である。図2には、電車の到来を利用者に案内する「まもなく1番線に電車が参ります」という音声を表す音響信号Sが例示されている。図2に例示される通り、音響信号Sは、「まもなく」という音片を表す音片信号Qと、「1番線に」という音片を表す音片信号Qと、「電車が参ります」という音片を表す音片信号Qとで構成される。各音片信号Qの時間長は相違し得る。 The storage device 14 of the first embodiment stores the acoustic signal S. As illustrated in FIG. 2, the acoustic signal S is a time-domain signal generated in advance by arranging a plurality of sound piece signals Q (Q 1 to Q 3 ) in time series and interconnecting them. . The sound piece signal Q is a signal representing a waveform of a partial sound (hereinafter referred to as “sound piece”) that constitutes a speech sound. One sound piece is a sound that is generated by segmental units such as words, phrases, and phrases. FIG. 2 exemplifies an acoustic signal S representing a voice “soon to arrive on the first line” that guides the user to the arrival of the train. As illustrated in FIG. 2, the sound signal S includes a sound piece signal Q 1 representing a sound piece “soon”, a sound piece signal Q 2 representing a sound piece “on line 1”, and “a train is coming. composed of the vibrating bar signal Q 3 representing the speech piece that ". The time length of each sound piece signal Q can be different.

以上の説明から理解される通り、例えば事前に録音された音片を表す複数の音片信号Qを接続する録音編集方式の音声合成技術により音響信号Sは事前に生成される。複数の音片信号Qの配列(総数,組合せ,順番)を変更することで、多様な発話内容を表す音響信号Sが生成される。音響信号Sを構成する複数の音片信号Qの配列は未知である。第1実施形態の音響解析装置100は、音響信号Sを構成する複数の音片信号Qの時系列を推定する。   As understood from the above description, for example, the acoustic signal S is generated in advance by a recording and editing type speech synthesis technique in which a plurality of sound piece signals Q representing sound pieces recorded in advance are connected. By changing the arrangement (total number, combination, order) of the plurality of sound piece signals Q, acoustic signals S representing various utterance contents are generated. The arrangement of the plurality of sound piece signals Q constituting the acoustic signal S is unknown. The acoustic analysis device 100 according to the first embodiment estimates a time series of a plurality of sound piece signals Q constituting the acoustic signal S.

第1実施形態の記憶装置14は、相異なる音片を表す複数(N個)の音片信号Qを記憶する。N個の音片信号Qの各々には相異なる番号(以下「音片番号」という)が付与される。音片番号n(n=1〜N)は、音片信号Qを識別するための識別情報である。記憶装置14に記憶されたN個の音片信号Qは、音響信号Sの解析に使用される。すなわち、第1実施形態の音響解析装置100は、記憶装置14に記憶されたN個の音片信号Qの各々と音響信号Sとを相互に対比することで、音響信号Sを構成する複数の音片信号Qの時系列を推定する。   The storage device 14 of the first embodiment stores a plurality (N) of sound piece signals Q representing different sound pieces. Each of the N sound piece signals Q is assigned a different number (hereinafter referred to as “sound piece number”). The sound piece number n (n = 1 to N) is identification information for identifying the sound piece signal Q. The N sound piece signals Q stored in the storage device 14 are used for the analysis of the acoustic signal S. In other words, the acoustic analysis device 100 according to the first embodiment compares each of the N sound piece signals Q stored in the storage device 14 with the acoustic signal S, so that a plurality of components constituting the acoustic signal S are compared. The time series of the sound piece signal Q is estimated.

制御装置12は、記憶装置14に記憶されたプログラムを実行することで、音響信号Sから複数の音片信号Qの時系列を推定するための解析処理部20として機能する。なお、制御装置12の機能を複数の装置に分散した構成、または、制御装置12の機能の少なくとも一部を専用の電子回路が実現する構成も採用され得る。   The control device 12 functions as an analysis processing unit 20 for estimating a time series of a plurality of sound piece signals Q from the acoustic signal S by executing a program stored in the storage device 14. A configuration in which the function of the control device 12 is distributed to a plurality of devices, or a configuration in which a dedicated electronic circuit realizes at least a part of the function of the control device 12 may be employed.

解析処理部20は、記憶装置14に記憶されたN個の音片信号Qの各々と音響信号Sとを対比することで、音響信号Sを構成する複数の音片信号Qの時系列を推定する。図1に例示される通り、第1実施形態の解析処理部20は、相関解析部22と推定処理部24とを含んで構成される。   The analysis processing unit 20 estimates the time series of the plurality of sound piece signals Q constituting the sound signal S by comparing each of the N sound piece signals Q stored in the storage device 14 with the sound signal S. To do. As illustrated in FIG. 1, the analysis processing unit 20 of the first embodiment includes a correlation analysis unit 22 and an estimation processing unit 24.

相関解析部22は、記憶装置14に記憶されたN個の音片信号Qの各々について、当該音片信号Qと音響信号Sとの相互相関Ct,nを単位期間(フレーム)毎に算定する。記号tは、音響信号Sを時間軸上で区分した複数(T個)の単位期間のうち任意の1個の単位期間を示す変数である(t=1〜T)。具体的には、音片番号nの音片信号Qと音響信号Sとの相互相関Ct,nは、以下の数式(1)で表現される。

Figure 2018146901
The correlation analysis unit 22 calculates the cross-correlation C t, n between the sound piece signal Q and the acoustic signal S for each of the N sound piece signals Q stored in the storage device 14 for each unit period (frame). To do. The symbol t is a variable indicating any one unit period among a plurality (T) of unit periods obtained by dividing the acoustic signal S on the time axis (t = 1 to T). Specifically, the cross-correlation C t, n between the sound piece signal Q of the sound piece number n and the acoustic signal S is expressed by the following formula (1).
Figure 2018146901

数式(1)の記号Xは、音響信号Sのうち第t番目の単位期間における周波数スペクトルである。また、数式(1)の記号Dn,tは、音片番号nの音片信号Qのうち第t番目の単位期間における周波数スペクトルである。周波数スペクトルXおよび周波数スペクトルDn,tの各々は、周波数軸上の相異なる周波数(周波数ビン)に対応する複数の数値の系列で表現され、例えば短時間フーリエ変換等の公知の周波数解析により算定される。 A symbol X t in Expression (1) is a frequency spectrum in the t-th unit period of the acoustic signal S. Symbol D n, t in Equation (1) is a frequency spectrum in the t-th unit period in the sound piece signal Q of the sound piece number n. Each of the frequency spectrum X t and the frequency spectrum D n, t is expressed by a series of a plurality of numerical values corresponding to different frequencies (frequency bins) on the frequency axis, and is obtained by a known frequency analysis such as a short-time Fourier transform. Calculated.

数式(1)の記号Lは、音片番号nの音片信号Qの時間長である。また、数式(1)の記号 ̄は複素共役を意味し、数式(1)の記号*は要素毎の積を意味する。数式(1)の記号F−1は、逆離散フーリエ変換である。 The symbol L n in the formula (1) is the time length of the sound piece signal Q of the sound piece number n. The symbol  ̄ in the formula (1) means complex conjugate, and the symbol * in the formula (1) means a product for each element. Symbol F −1 in Equation (1) is an inverse discrete Fourier transform.

以上の説明から理解される通り、音響信号Sに対する音片番号nの音片信号Qの時間軸上の位置を変化させた場合に、音片信号Qと音響信号Sとの間で波形が類似するほど、当該音片信号Qの末尾に相当する時点tの相互相関Ct,nは大きい数値となる。すなわち、図3に例示される通り、音響信号Sのうち相互相関Ct,nが極大となる時点tに末尾が一致するように音片信号Qを配置した状態で、音響信号Sの波形と音片信号Qの波形とが類似する。音響信号Sのうち相互相関Ct,nが極大となる時点tから逆方向(時間を遡及する方向)の時間長Lにわたる区間が、音片番号nの音片信号Qの波形に類似する、とも換言され得る。以上の説明から理解される通り、第1実施形態の相関解析部22は、N個の音片信号Qの各々と音響信号Sとを対比する要素として表現される。 As understood from the above description, when the position on the time axis of the sound piece signal Q of the sound piece number n with respect to the sound signal S is changed, the waveform is similar between the sound piece signal Q and the sound signal S. The cross-correlation C t, n at the time point t corresponding to the end of the sound piece signal Q becomes a larger numerical value. That is, as illustrated in FIG. 3, in the state where the sound piece signal Q is arranged so that the end coincides with the time point t at which the cross-correlation C t, n of the sound signal S becomes maximum, The waveform of the sound piece signal Q is similar. In the acoustic signal S, a section extending from the time point t at which the cross-correlation C t, n is maximized to the time length L n in the reverse direction (the direction in which the time is retroactive) is similar to the waveform of the sound piece signal Q of the sound piece number n. In other words. As understood from the above description, the correlation analysis unit 22 of the first embodiment is expressed as an element that compares each of the N sound piece signals Q with the acoustic signal S.

図4に例示される通り、時間軸上で相互に接続された複数の音片信号Qの時系列(以下「音片系列」という)Zを想定する。図4には、3個の音片信号Q〜Qのうちの2個の組合せで構成された2通りの音片系列Z(Z12,Z23)が便宜的に図示されている。また、3個の音片信号Q〜Qの各々について、当該音片信号Qと音響信号Sとの間で算定された相互相関Ct,n(Ct,1,Ct,2,Ct,3)が図4には併記されている。 As illustrated in FIG. 4, a time series (hereinafter referred to as “speech series”) Z of a plurality of sound piece signals Q connected to each other on the time axis is assumed. In FIG. 4, two sound piece sequences Z (Z 12 , Z 23 ) composed of a combination of two of the three sound piece signals Q 1 to Q 3 are shown for convenience. For each of the three sound piece signals Q 1 to Q 3 , the cross-correlation C t, n (C t, 1 , C t, 2 , C t, 3 ) is also shown in FIG.

いま、図4に例示される通り、音片系列Zを構成する各音片信号Qの末尾の時点における当該音片信号Qと音響信号Sとの相互相関Ct,nを、音片系列Zを構成する複数の音片信号Qについて累積した数値(以下「累積相互相関」という)Rを検討する。 Now, as illustrated in FIG. 4, the cross-correlation C t, n between the sound piece signal Q and the acoustic signal S at the end of each sound piece signal Q constituting the sound piece sequence Z is expressed as the sound piece sequence Z. A numerical value (hereinafter referred to as “cumulative cross-correlation”) R that is accumulated for a plurality of sound piece signals Q that constitutes is considered.

例えば、音片信号Qに音片信号Qを後続させた音片系列Z12については、相互相関Cta,1と相互相関Ctb,2との合計値が累積相互相関R12として算定される。相互相関Cta,1は、音片信号Qと音響信号Sとの相互相関C1,1〜CT,1のうち、音片系列Z12における音片信号Qの末尾の時点tに対応する数値である。他方、相互相関Ctb,2は、音片信号Qと音響信号Sとの相互相関C1,2〜CT,2のうち、音片系列Z12における音片信号Qの末尾の時点tに対応する数値である。 For example, for the sound piece sequence Z 12 in which the sound piece signal Q 1 is followed by the sound piece signal Q 2 , the total value of the cross-correlation C ta, 1 and the cross-correlation C tb, 2 is calculated as the cumulative cross-correlation R 12. Is done. The cross-correlation C ta, 1 is a time t a at the end of the sound piece signal Q 1 in the sound piece sequence Z 12 among the cross-correlations C 1,1 to C T, 1 between the sound piece signal Q 1 and the acoustic signal S. It is a numerical value corresponding to. On the other hand, the cross-correlation C tb, 2, the sound piece signal Q 2 and of the cross-correlation C 1, 2 -C T, 2 the acoustic signal S, the speech piece signal Q 2 at the speech segment sequence Z 12 end point is a numerical value that corresponds to t b.

また、音片信号Qに音片信号Qを後続させた音片系列Z23については、相互相関Ctc,2と相互相関Ctd,3との合計値が累積相互相関R23として算定される。相互相関Ctc,2は、音片信号Qと音響信号Sとの相互相関C1,2〜CT,2のうち音片系列Z23における音片信号Qの末尾の時点tに対応する数値である。相互相関Ctd,3は、音片信号Qと音響信号Sとの相互相関C1,3〜CT,3のうち音片系列Z23における音片信号Qの末尾の時点tに対応する数値である。 As for the speech segment sequence Z 23 in which subsequent to cause the tuning bar signal Q 3 in the speech segment signal Q 2, calculating the total value of the cross-correlation C tc, 2 and the cross-correlation C td, 3 as the cumulative cross-correlation R 23 Is done. The cross-correlation C tc, 2 is at the time t c at the end of the sound piece signal Q 2 in the sound piece sequence Z 23 among the cross-correlations C 1,2 to C T, 2 between the sound piece signal Q 2 and the acoustic signal S. Corresponding numerical value. The cross-correlation C td, 3 is the time t d at the end of the sound piece signal Q 3 in the sound piece sequence Z 23 among the cross-correlations C 1,3 to C T, 3 between the sound piece signal Q 3 and the acoustic signal S. Corresponding numerical value.

図4では、音響信号Sが実際には音片信号Qと音片信号Qとで構成される場合(すなわち音片系列Z12が正解である場合)が想定されている。したがって、音片系列Z12における音片信号Qの末尾の時点tにおける相互相関Cta,1と音片信号Qの末尾の時点tにおける相互相関Ctb,2とは大きい数値(最大値1に近い数値)となる。すなわち、累積相互相関R12は大きい数値となる。他方、音片系列Z23における音片信号Qの末尾の時点tにおける相互相関Ctc,2と音片信号Qの末尾の時点tにおける相互相関Ctd,3とは小さい数値となる。すなわち、累積相互相関R23は小さい数値となる。 In FIG. 4, it is assumed that the acoustic signal S is actually composed of the sound piece signal Q 1 and the sound piece signal Q 2 (that is, the sound piece sequence Z 12 is correct). Therefore, the cross-correlation C ta, 1 sound piece signal cross-correlation C tb at the end of the time t b of Q 2, 2 larger numbers at the end of the time t a of the speech piece signal Q 1 at the speech piece series Z 12 ( (A value close to the maximum value 1). That is, the cumulative cross-correlation R 12 becomes large numbers. On the other hand, the cross-correlation C tc, 2 at the end time t c of the sound piece signal Q 2 in the sound piece sequence Z 23 and the cross-correlation C td, 3 at the end time t d of the sound piece signal Q 3 are small numerical values. Become. That is, the cumulative cross-correlation R 23 is a smaller number.

以上の説明から理解される通り、音響信号Sを構成する複数の音片信号Qの組合せに音片系列Zが近いほど、当該音片系列Zについて算定される累積相互相関Rは大きい数値となる。したがって、音片系列Zの適否を評価するための指標として累積相互相関Rを利用可能である。すなわち、音片系列Zの累積相互相関Rが大きいほど、音響信号Sを構成する複数の音片信号Qの組合せとして当該音片系列Zが適正であると評価できる。以上の傾向を背景として、推定処理部24は、累積相互相関Rが最大化されるように音片系列Z(複数の音片信号Qの時系列)を推定する。   As understood from the above description, the closer the sound piece sequence Z is to the combination of the plurality of sound piece signals Q constituting the acoustic signal S, the larger the cumulative cross-correlation R calculated for the sound piece sequence Z becomes. . Therefore, the cumulative cross-correlation R can be used as an index for evaluating the suitability of the sound piece series Z. That is, as the cumulative cross-correlation R of the sound piece sequence Z is larger, it can be evaluated that the sound piece sequence Z is more appropriate as a combination of a plurality of sound piece signals Q constituting the acoustic signal S. Against the background described above, the estimation processing unit 24 estimates the sound piece series Z (the time series of the plurality of sound piece signals Q) so that the cumulative cross-correlation R is maximized.

ところで、累積相互相関Rを最大化する音片系列Zを推定する方法としては、複数の音片信号Qを配列する全通りの順列(音片系列Z)について累積相互相関Rを算定し、累積相互相関Rが最大となる音片系列Zを選択する方法も想定される。しかし、以上の方法では演算量が膨大となる可能性がある。そこで、第1実施形態の推定処理部24は、推定処理を効率化し得る動的計画法を利用して、累積相互相関Rを最大化する音片系列Zを探索する。推定処理部24の具体的な動作を以下に詳述する。   By the way, as a method of estimating the sound piece sequence Z that maximizes the cumulative cross-correlation R, the cumulative cross-correlation R is calculated for all permutations (speech sequence Z) in which a plurality of sound piece signals Q are arranged. A method of selecting a sound piece sequence Z that maximizes the cross-correlation R is also assumed. However, the amount of calculation may be enormous in the above method. Therefore, the estimation processing unit 24 according to the first embodiment searches for the sound piece series Z that maximizes the cumulative cross-correlation R by using dynamic programming that can make the estimation process more efficient. The specific operation of the estimation processing unit 24 will be described in detail below.

音響信号Sの時点tに音片番号nの音片信号Qの末尾が位置すると仮定すると、当該時点tにおける累積相互相関Rt,nは、以下の数式(2)で表現される。数式(2)の累積相互相関Rt,nは、N個の音片信号Qの各々について算定される。

Figure 2018146901

数式(2)のうち右辺の第1項は、時点tから音片番号nの音片信号Qの時間長Lだけ遡及した時点(t−L)について算定されたN個の累積相互相関Rt−Ln,1〜Rt−Ln,Nの最大値(max)である。推定処理部24は、現在の時点tについて相関解析部22が算定した相互相関Ct,nを当該最大値に加算することで、累積相互相関Rt,nを算定する。 Assuming that the end of the sound piece signal Q having the sound piece number n is located at the time t of the acoustic signal S, the cumulative cross-correlation R t, n at the time t is expressed by the following equation (2). The cumulative cross-correlation R t, n in Equation (2) is calculated for each of the N sound piece signals Q.
Figure 2018146901

The first term on the right side of Equation (2) is the N cumulative cross-correlations calculated for the time point (t−L n ) retroactive from the time point t by the time length L n of the sound piece signal Q of the sound piece number n. It is the maximum value (max) of R t-Ln, 1 to R t-Ln, N. The estimation processing unit 24 calculates the cumulative cross-correlation R t, n by adding the cross-correlation C t, n calculated by the correlation analysis unit 22 for the current time t to the maximum value.

また、数式(2)におけるRt−Ln,mを最大化させる音片信号Qの音片番号It,nは、以下の数式(3)で表現される。

Figure 2018146901
Further, the sound piece number It , n of the sound piece signal Q that maximizes R t−Ln, m in Expression (2) is expressed by Expression (3) below.
Figure 2018146901

各音片番号nについて音響信号Sの末尾までの累積相互相関R1,n〜RT,nを算定すると、推定処理部24は、音響信号Sの末尾(t=t=T)におけるN個の累積相互相関RT,1〜RT,Nの最大値に対応する音片信号Qの音片番号nを選択する。すなわち、音片番号nは以下の数式(4)で表現される。

Figure 2018146901
When the cumulative cross-correlation R 1, n to R T, n up to the end of the acoustic signal S is calculated for each sound piece number n, the estimation processing unit 24 determines N at the end (t = t 0 = T) of the acoustic signal S. A sound piece number n 0 of the sound piece signal Q corresponding to the maximum value of the cumulative cross-correlations RT, 1 to RT, N is selected. That is, the sound piece number n 0 is expressed by the following mathematical formula (4).
Figure 2018146901

数式(4)の音片番号nは、音響信号Sの最後に位置する音片信号Qの音片番号nである。音片番号nの音片信号Qの直前に位置すべき音片信号Qの音片番号nは、数式(3)から理解される通り、音片番号It0,n0であり、当該直前の音片信号Qの末尾の時点tは、時点tから第n番目の音片信号Qの時間長Ln0だけ遡及した時点(t−Ln0)である。以上の説明から理解される通り、音響信号Sの末尾から逆方向に音片信号Qを辿る処理(バックトラック)は、以下の数式(5)および数式(6)の漸化式で表現される。

Figure 2018146901

Figure 2018146901
The sound piece number n 0 in Equation (4) is the sound piece number n of the sound piece signal Q located at the end of the acoustic signal S. Speech segment number n 1 of the speech piece signal Q to be located just before the speech piece signal Q of speech segment number n 0, as will be understood from the formula (3) is a speech segment number I t0, n0, the immediately preceding time t 1 the end of the vibrating bar signal Q is the time of the retroactively from the time t 0 by the time length L n0 of the n 0 th speech unit signal Q (t 0 -L n0). As understood from the above description, the process (backtrack) of tracing the sound piece signal Q in the reverse direction from the end of the acoustic signal S is expressed by the recurrence formulas of the following formulas (5) and (6). .
Figure 2018146901

Figure 2018146901

推定処理部24は、数式(5)および数式(6)で表現されるバックトラックを、音響信号Sの始点に到達するまで反復する。以上の手順で探索した音片番号nの系列{n,nI−1,…,n,n}により、音響信号Sを構成する複数の音片信号Qの時系列(音片系列Z)が表現される。 The estimation processing unit 24 repeats the backtrack expressed by Expression (5) and Expression (6) until the start point of the acoustic signal S is reached. The time series (speech piece) of the plurality of sound piece signals Q constituting the acoustic signal S by the series {n I , n I−1 ,..., N 1 , n 0 } of the sound piece numbers n i searched in the above procedure. A sequence Z) is represented.

図5は、第1実施形態の制御装置12が音片系列Zを推定する動作(以下「音響解析処理」という)のフローチャートである。例えば利用者からの指示を契機として音響解析処理が開始される。音響解析処理を開始すると、相関解析部22は、記憶装置14に記憶されたN個の音片信号Qの各々について、当該音片信号Qと音響信号Sとの相互相関Ct,nを単位期間(フレーム)毎に算定する(S1)。相互相関Ct,nの算定が完了すると、推定処理部24は、累積相互相関Rを最大化する音片系列Zを、前述の動的計画法により推定する(S2)。 FIG. 5 is a flowchart of an operation (hereinafter referred to as “acoustic analysis process”) in which the control device 12 of the first embodiment estimates the sound piece series Z. For example, the acoustic analysis process is started in response to an instruction from the user. When the acoustic analysis process is started, the correlation analysis unit 22 uses the cross-correlation C t, n between the sound piece signal Q and the sound signal S as a unit for each of the N sound piece signals Q stored in the storage device 14. Calculation is performed for each period (frame) (S1). When the calculation of the cross-correlation C t, n is completed, the estimation processing unit 24 estimates the sound piece sequence Z that maximizes the cumulative cross-correlation R by the above-described dynamic programming (S2).

以上に説明した通り、第1実施形態では、N個の音片信号Qの各々と音響信号Sとを対比することで、音響信号Sを構成する複数の音片信号Qの時系列を推定することが可能である。第1実施形態では特に、各音片信号Qの末尾の時点tにおける相互相関Ct,nを複数の音片信号Qの時系列にわたり累積した数値(累積相互相関R)が最大化されるように、複数の音片信号Qの時系列(音片系列Z)が推定される。したがって、各音片信号Qと音響信号Sとの波形の類似性という観点から、音響信号Sを構成する複数の音片信号Qの時系列を高精度に推定できるという利点がある。また、第1実施形態では、動的計画法により音片系列Zが推定される。したがって、例えば複数の音片信号Qを配列する全通りの順列について累積相互相関Rを算定したうえで、累積相互相関Rが最大となる音片系列Zを選択する方法と比較して、制御装置12による演算量を削減することが可能である。ただし、複数の音片信号Qの全通りの順列について累積相互相関Rを算定したうえで最大値を探索する方法を採用してもよい。 As described above, in the first embodiment, the time series of the plurality of sound piece signals Q constituting the sound signal S is estimated by comparing each of the N sound piece signals Q with the sound signal S. It is possible. In particular, in the first embodiment, the numerical value (cumulative cross-correlation R) obtained by accumulating the cross-correlation C t, n at the end time t of each sound piece signal Q over the time series of the plurality of sound piece signals Q is maximized. In addition, a time series (speech series Z) of a plurality of sound piece signals Q is estimated. Therefore, there is an advantage that the time series of the plurality of sound piece signals Q constituting the sound signal S can be estimated with high accuracy from the viewpoint of the waveform similarity between each sound piece signal Q and the sound signal S. In the first embodiment, the sound piece sequence Z is estimated by dynamic programming. Therefore, for example, after calculating the cumulative cross-correlation R for all permutations in which a plurality of sound piece signals Q are arranged, the control device is compared with a method of selecting a sound piece sequence Z that maximizes the cumulative cross-correlation R. The amount of computation by 12 can be reduced. However, a method of searching for the maximum value after calculating the cumulative cross-correlation R for all the permutations of the plurality of sound piece signals Q may be employed.

<第2実施形態>
本発明の第2実施形態について説明する。なお、以下に例示する各形態において作用または機能が第1実施形態と同様である要素については、第1実施形態の説明で使用した符号を流用して各々の詳細な説明を適宜に省略する。
Second Embodiment
A second embodiment of the present invention will be described. In addition, about the element which an effect | action or function is the same as that of 1st Embodiment in each form illustrated below, the code | symbol used by description of 1st Embodiment is diverted, and each detailed description is abbreviate | omitted suitably.

第1実施形態では、複数の音片信号Qが時間軸上に重複も隙間もなく配列されることで音響信号Sが構成されるから、相前後する2個の音片信号Qの末尾の間隔(時点ti−1と時点tとの間隔)は音片信号Qの時間長Lに一致する。しかし、実際には、相前後する2個の音片信号Qが相互に重複または離間した状態で配列され得る。すなわち、相前後する2個の音片信号Qの末尾の間隔(時点ti−1と時点tとの間隔)は、音片信号Qの時間長Lとは僅かに相違した時間長である可能性がある。以上の事情を考慮して、第2実施形態では、相前後する2個の音片信号Qの末尾の間隔に誤差εを加味したうえで、音響信号Sを構成する複数の音片信号Qの時系列(音片系列Z)を推定する。 In the first embodiment, since the sound signal S is formed by arranging a plurality of sound piece signals Q on the time axis without overlapping or gaps, the interval between the last two sound piece signals Q ( The interval between the time point t i-1 and the time point t i is equal to the time length L n of the sound piece signal Q. However, in practice, two adjacent sound piece signals Q may be arranged in a state where they overlap or are separated from each other. In other words, the end interval between the two adjacent sound piece signals Q (interval between the time point t i-1 and the time point t i ) is a time length slightly different from the time length L n of the sound piece signal Q. There is a possibility. In consideration of the above circumstances, in the second embodiment, the error ε is added to the end interval between two adjacent sound piece signals Q, and a plurality of sound piece signals Q constituting the sound signal S are added. A time series (speech piece series Z) is estimated.

具体的には、第2実施形態では、第1実施形態の数式(2)が以下の数式(2a)に置換される。

Figure 2018146901

すなわち、数式(2a)における右辺の第1項は、時点tから音片番号nの音片信号Qの時間長Lだけ遡及した時点(t−L)を中心として幅2Eの範囲内(t−L±E)におけるN個の累積相互相関Rt−Ln+ε,1〜Rt−Ln+ε,Nの最大値(max)である。定数Eは所定の正数である。数式(2a)における累積相互相関Rt−Ln+ε,mを最大化させる音片信号Qの音片番号It,nは、以下の数式(3a)で表現される。すなわち、第1実施形態の数式(3)が第2実施形態では数式(3a)に置換される。
Figure 2018146901
Specifically, in the second embodiment, the formula (2) of the first embodiment is replaced with the following formula (2a).
Figure 2018146901

In other words, the first term on the right side in the formula (2a) is within the range of the width 2E around the time point (t−L n ) retroactive by the time length L n of the sound piece signal Q of the sound piece number n from the time point t ( This is the maximum value (max) of N cumulative cross-correlations R t−Ln + ε, 1 to R t−Ln + ε, N at t−L n ± E). The constant E is a predetermined positive number. The sound piece number I t, n of the sound piece signal Q that maximizes the cumulative cross-correlation R t−Ln + ε, m in Expression (2a) is expressed by Expression (3a) below. That is, Formula (3) in the first embodiment is replaced with Formula (3a) in the second embodiment.
Figure 2018146901

また、累積相互相関Rt−Ln+ε,mを最大化させる誤差εは、以下の数式(7)で表現される。

Figure 2018146901
Further, the error ε that maximizes the cumulative cross-correlation R t−Ln + ε, m is expressed by the following equation (7).
Figure 2018146901

第2実施形態の推定処理部24が実行するバックトラックは、第1実施形態の数式(6)に数式(7)の誤差Jt,nを導入した以下の数式(6a)で表現される。

Figure 2018146901

数式(6a)から理解される通り、推定処理部24は、時点ti−1に対して音片信号Qの時間長Lni−1だけ遡及した時点から更に誤差Jti−1,ni−1だけずれた時点を時間軸上の逆方向に辿る。すなわち、推定処理部24は、第1実施形態と同様の数式(5)と誤差Jt,nを含む数式(6a)とで表現されるバックトラックにより、音片番号nの系列{n,nI−1,…,n,n}(音響信号Sを構成する複数の音片信号Qの時系列)を推定する。 The backtrack executed by the estimation processing unit 24 of the second embodiment is expressed by the following formula (6a) in which the error J t, n of the formula (7) is introduced into the formula (6) of the first embodiment.
Figure 2018146901

As understood from the mathematical expression (6a), the estimation processing unit 24 further increases the error J ti−1, ni−1 from the time point L ni−1 retroactive to the time point t i−1 . Traces the time point shifted in the opposite direction on the time axis. That is, the estimation processing unit 24, the same formula as in the first embodiment (5) the error J t, the backtracking expressed out with equations (6a) containing n, speech segment number n i of sequence {n I , N I−1 ,..., N 1 , n 0 } (time series of a plurality of sound piece signals Q constituting the acoustic signal S) is estimated.

第2実施形態においても第1実施形態と同様の効果が実現される。また、第2実施形態では、各音片信号Qの時間長Lに誤差εが加味されるから、相前後する2個の音片信号Qが相互に重複または離間している場合でも、音響信号Sを構成する複数の音片信号Qの時系列を高精度に推定できるという利点がある。 In the second embodiment, the same effect as in the first embodiment is realized. In the second embodiment, since the error ε is added to the time length L n of each sound piece signal Q, even if two adjacent sound piece signals Q overlap or are separated from each other, the sound There is an advantage that the time series of the plurality of sound piece signals Q constituting the signal S can be estimated with high accuracy.

<第3実施形態>
図6は、第1実施形態または第2実施形態の音響解析装置100を利用した情報提供装置200の構成図である。図6に例示される通り、第3実施形態の情報提供装置200は、制御装置32と記憶装置34と収音装置36と放音装置38とを具備するコンピュータシステムで実現される。なお、情報提供装置200は、単体の装置として実現されるほか、相互に別体で構成された複数の装置でも実現され得る。
<Third Embodiment>
FIG. 6 is a configuration diagram of an information providing apparatus 200 using the acoustic analysis apparatus 100 of the first embodiment or the second embodiment. As illustrated in FIG. 6, the information providing apparatus 200 according to the third embodiment is realized by a computer system including a control device 32, a storage device 34, a sound collection device 36, and a sound emission device 38. Note that the information providing apparatus 200 can be realized as a single apparatus or a plurality of apparatuses configured separately from each other.

制御装置32は、例えばCPU等の処理回路で構成され、情報提供装置200の動作を統括的に制御する。記憶装置34は、制御装置32が実行するプログラムと制御装置32が使用する各種のデータとを記憶する。例えば磁気記録媒体および半導体記録媒体等の公知の記録媒体が記憶装置34として利用され得る。第3実施形態の記憶装置34は、前述の各形態で例示したN個の音片信号Qを記憶する。   The control device 32 is configured by a processing circuit such as a CPU, for example, and comprehensively controls the operation of the information providing device 200. The storage device 34 stores a program executed by the control device 32 and various data used by the control device 32. For example, a known recording medium such as a magnetic recording medium and a semiconductor recording medium can be used as the storage device 34. The storage device 34 of the third embodiment stores N sound piece signals Q exemplified in the above-described embodiments.

収音装置36は、交通施設または商業施設等の各種の施設で発音または放送された案内用の音声(以下「案内音声」という)Gを収音することで、当該案内音声Gを表す音響信号Sを生成する。音響信号Sが表す案内音声Gは、複数の音片の時系列である。放音装置38は、制御装置32による制御のもとで音を再生する。   The sound collection device 36 collects a guidance voice (hereinafter referred to as “guidance voice”) G that is sounded or broadcasted at various facilities such as a traffic facility or a commercial facility, so that an acoustic signal representing the guidance voice G is collected. S is generated. The guidance voice G represented by the acoustic signal S is a time series of a plurality of sound pieces. The sound emitting device 38 reproduces sound under the control of the control device 32.

制御装置32は、記憶装置34に記憶されたプログラムを実行することで、第1実施形態または第2実施形態で例示した解析処理部20に加えて、変調処理部42および混合処理部44として機能する。なお、制御装置32の機能を複数の装置に分散した構成、または、制御装置32の機能の少なくとも一部を専用の電子回路が実現する構成も採用され得る。解析処理部20は、第1実施形態または第2実施形態で例示した構成および動作により、収音装置36が生成した音響信号Sから音片系列Zを推定する。すなわち、第3実施形態の情報提供装置200は、第1実施形態または第2実施形態の音響解析装置100を含んで構成される。したがって、第3実施形態においても第1実施形態または第2実施形態と同様の効果が実現される。   The control device 32 functions as a modulation processing unit 42 and a mixing processing unit 44 in addition to the analysis processing unit 20 exemplified in the first embodiment or the second embodiment by executing a program stored in the storage device 34. To do. A configuration in which the function of the control device 32 is distributed to a plurality of devices or a configuration in which a dedicated electronic circuit realizes at least a part of the function of the control device 32 may be employed. The analysis processing unit 20 estimates the sound piece sequence Z from the acoustic signal S generated by the sound collection device 36 by the configuration and operation exemplified in the first embodiment or the second embodiment. That is, the information providing apparatus 200 of the third embodiment is configured to include the acoustic analysis apparatus 100 of the first embodiment or the second embodiment. Therefore, also in the third embodiment, the same effect as the first embodiment or the second embodiment is realized.

変調処理部42は、解析処理部20が推定した音片系列Zに応じた変調信号Mを生成する。変調信号Mは、音片系列Zに応じた配信情報Bを音響成分として含む信号である。配信情報Bは、例えば音片系列Z自体または当該音片系列Zを識別するための識別情報である。変調処理部42は、例えば所定の周波数の正弦波等の搬送波を配信情報Bにより変調する周波数変調、または、拡散符号を利用した配信情報Bの拡散変調等の変調処理により変調信号Mを生成する。配信情報Bの音響成分の周波数帯域は、例えば、放音装置38による再生が可能な周波数帯域であり、かつ、利用者が通常の環境で聴取する音の周波数帯域を上回る範囲(例えば18kHz以上かつ20kHz以下)に包含される。   The modulation processing unit 42 generates a modulation signal M corresponding to the sound piece sequence Z estimated by the analysis processing unit 20. The modulation signal M is a signal including distribution information B corresponding to the sound piece series Z as an acoustic component. The distribution information B is identification information for identifying the sound piece sequence Z itself or the sound piece sequence Z, for example. The modulation processing unit 42 generates the modulation signal M by modulation processing such as frequency modulation for modulating a carrier wave such as a sine wave of a predetermined frequency with the distribution information B, or spread modulation of the distribution information B using a spread code. . The frequency band of the acoustic component of the distribution information B is, for example, a frequency band that can be reproduced by the sound emitting device 38, and a range that exceeds the frequency band of the sound that the user listens to in a normal environment (for example, 18 kHz or more and 20 kHz or less).

混合処理部44は、収音装置36から供給される音響信号Sと変調処理部42が生成した変調信号Mとを混合(例えば加算)することで音響信号Yを生成する。放音装置38は、音響信号Yが表す音を放音する。すなわち、音響信号Sが表す案内音声Gと変調信号Mが表す配信情報Bの音響成分とが放音装置38から再生される。以上の説明から理解される通り、第1実施形態の放音装置38は、案内音声Gを再生する音響機器として機能するほか、空気振動としての音波を伝送媒体とした音響通信で配信情報Bを送信する送信機としても機能する。   The mixing processing unit 44 generates the acoustic signal Y by mixing (for example, adding) the acoustic signal S supplied from the sound collection device 36 and the modulation signal M generated by the modulation processing unit 42. The sound emitting device 38 emits the sound represented by the acoustic signal Y. That is, the guidance sound G represented by the acoustic signal S and the acoustic component of the distribution information B represented by the modulation signal M are reproduced from the sound emitting device 38. As understood from the above description, the sound emitting device 38 of the first embodiment functions as an acoustic device that reproduces the guidance voice G, and also distributes the distribution information B by acoustic communication using sound waves as air vibration as a transmission medium. It also functions as a transmitter to transmit.

図6の端末装置300は、例えば携帯電話機またはスマートフォン等の情報端末である。なお、例えば、電光掲示板または電子看板(例えばデジタルサイネージ)等の案内用の表示端末を端末装置300として利用することも可能である。第3実施形態の端末装置300は、情報提供装置200による再生音から配信情報Bを復調し、当該配信情報Bに対応する関連情報を出力装置(例えば表示装置または放音装置)から出力する。案内音声Gの音響信号Sから生成された配信情報Bが示す関連情報は、当該案内音声Gに関連する情報(例えば案内音声Gを表す文字列やその翻訳文)である。以上の説明から理解される通り、端末装置300の利用者は、情報提供装置200が再生する案内音声Gを聴取するほか、当該案内音声Gに対応する関連情報を端末装置300により確認することが可能である。   6 is an information terminal such as a mobile phone or a smartphone. Note that, for example, a display terminal for guidance such as an electronic bulletin board or an electronic signboard (for example, digital signage) can be used as the terminal device 300. The terminal device 300 according to the third embodiment demodulates the distribution information B from the reproduced sound by the information providing device 200, and outputs related information corresponding to the distribution information B from an output device (for example, a display device or a sound emitting device). The related information indicated by the distribution information B generated from the acoustic signal S of the guidance voice G is information related to the guidance voice G (for example, a character string representing the guidance voice G or a translation thereof). As understood from the above description, the user of the terminal device 300 can listen to the guidance voice G reproduced by the information providing apparatus 200 and can confirm related information corresponding to the guidance voice G by the terminal device 300. Is possible.

なお、以上の説明では、配信情報Bを音響通信により端末装置300に送信したが、配信情報Bを送信するための通信方式は以上の例示に限定されない。例えば、電磁波を伝送媒体として利用した無線通信(典型的には近距離無線通信)により配信情報Bを端末装置300に送信することも可能である。また、関連情報を配信情報Bとして端末装置300に送信することも可能である。   In the above description, the distribution information B is transmitted to the terminal device 300 by acoustic communication. However, the communication method for transmitting the distribution information B is not limited to the above examples. For example, the distribution information B can be transmitted to the terminal device 300 by wireless communication (typically short-range wireless communication) using electromagnetic waves as a transmission medium. It is also possible to transmit the related information as the distribution information B to the terminal device 300.

<変形例>
以上に例示した各態様は多様に変形され得る。具体的な変形の態様を以下に例示する。以下の例示から任意に選択された2個以上の態様は、相互に矛盾しない範囲で適宜に併合され得る。
<Modification>
Each aspect illustrated above can be variously modified. Specific modifications are exemplified below. Two or more modes arbitrarily selected from the following examples can be appropriately combined within a range that does not contradict each other.

(1)音響信号Sのサンプル値を所定の比率で間引いたうえで各音片信号Qと対比してもよい。以上の構成によれば、制御装置12の演算量を削減することが可能である。 (1) The sample value of the acoustic signal S may be thinned out at a predetermined ratio and then compared with each sound piece signal Q. According to the above configuration, the calculation amount of the control device 12 can be reduced.

(2)数式(3)または数式(3a)の音片番号It,nと数式(7)の誤差Jt,nとは、音響信号Sの全区間にわたり保持する必要があるものの、推定処理部24が時点tの累積相互相関Rt,nを算定する段階では、時点(t−Lmax−ε)よりも過去の累積相互相関Rは不要である。時間長Lmaxは、N個の音片信号Qの時間長Lの最大値である。以上の説明から理解される通り、任意の時点tでは、{(Lmax+E)×N}個の累積相互相関Rを記憶装置14に保持すれば足りる。 (2) The sound piece number It , n in Equation (3) or Equation (3a) and the error J t, n in Equation (7) need to be held over the entire interval of the acoustic signal S, but are estimated processing At the stage where the unit 24 calculates the cumulative cross-correlation R t, n at the time point t, the past cumulative cross-correlation R from the time point (t−L max −ε) is unnecessary. The time length L max is the maximum value of the time length L n of the N sound piece signals Q. As understood from the above description, it is sufficient to store {(L max + E) × N} accumulated cross-correlations R in the storage device 14 at an arbitrary time point t.

(3)前述の各形態では、音声を表す音響信号Sを例示したが、音声以外の音(例えば楽音)を表す音響信号Sについても、前述の各形態と同様の方法により、当該音響信号Sを構成する複数の音片信号Qの時系列を推定することが可能である。したがって、各音片信号Qが表す音も音声には限定されない。例えば、相異なる楽曲から抽出された音片を表す複数の音片信号Qで構成される音響信号Sを解析することで、音響信号Sの素材として利用された音片(さらには楽曲名)を特定することが可能である。 (3) In each of the above-described embodiments, the acoustic signal S representing the voice is exemplified, but the acoustic signal S representing the sound other than the voice (for example, a musical sound) is also processed by the same method as in each of the above-described embodiments. It is possible to estimate a time series of a plurality of sound piece signals Q that constitute. Therefore, the sound represented by each sound piece signal Q is not limited to sound. For example, by analyzing the sound signal S composed of a plurality of sound piece signals Q representing sound pieces extracted from different music pieces, the sound pieces (and music names) used as the material of the sound signal S are obtained. It is possible to specify.

(4)第1実施形態および第2実施形態では、記憶装置14に記憶された音響信号Sを解析したが、第3実施形態での例示からも理解される通り、収音装置による収音で生成された音響信号Sを解析することも可能である。 (4) In the first embodiment and the second embodiment, the acoustic signal S stored in the storage device 14 is analyzed. However, as understood from the illustration in the third embodiment, the sound collection device collects sound. It is also possible to analyze the generated acoustic signal S.

(5)前述の各形態に係る音響解析装置100は、各形態での例示の通り、制御装置12とプログラムとの協働により実現される。前述の各形態に係るプログラムは、制御装置12(コンピュータの例示)に、N個の音片信号Qの各々と音響信号Sとを対比することで、音響信号Sを構成する複数の音片信号Qの時系列(音片系列Z)を推定する音響解析処理を実行させる。 (5) The acoustic analysis device 100 according to each embodiment described above is realized by the cooperation of the control device 12 and the program as illustrated in each embodiment. The program according to each of the above-described embodiments is obtained by comparing each of the N sound piece signals Q and the sound signal S with the control device 12 (exemplification of a computer), thereby a plurality of sound piece signals constituting the sound signal S. An acoustic analysis process for estimating a time series of Q (sound piece series Z) is executed.

以上に例示したプログラムは、コンピュータが読取可能な記録媒体に格納された形態で提供されてコンピュータにインストールされ得る。記録媒体は、例えば非一過性(non-transitory)の記録媒体であり、CD-ROM等の光学式記録媒体(光ディスク)が好例であるが、半導体記録媒体または磁気記録媒体等の公知の任意の形式の記録媒体を包含し得る。なお、非一過性の記録媒体とは、一過性の伝搬信号(transitory, propagating signal)を除く任意の記録媒体を含み、揮発性の記録媒体を除外するものではない。また、通信網を介した配信の形態でプログラムをコンピュータに提供することも可能である。   The programs exemplified above can be provided in a form stored in a computer-readable recording medium and installed in the computer. The recording medium is, for example, a non-transitory recording medium, and an optical recording medium (optical disk) such as a CD-ROM is a good example, but a known arbitrary one such as a semiconductor recording medium or a magnetic recording medium This type of recording medium can be included. Note that the non-transitory recording medium includes any recording medium except for a transient propagation signal (transitory, propagating signal), and does not exclude a volatile recording medium. It is also possible to provide a program to a computer in the form of distribution via a communication network.

(6)以上に例示した形態から、例えば以下の構成が把握される。
<態様1>
本発明の好適な態様(態様1)に係る音響解析方法は、N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する。以上の構成によれば、N個の音片信号の各々と音響信号とを対比することで、音響信号を構成する複数の音片信号の時系列を推定することが可能である。
<態様2>
態様1の好適例(態様2)において、前記複数の音片信号の時系列の推定は、前記N個の音片信号の各々について、当該音片信号と前記音響信号との相互相関を算定する相関解析と、前記N個の音片信号から選択した前記複数の音片信号の時系列を推定する推定処理とを含み、前記推定処理においては、時間軸上に配列された2以上の音片信号の末尾の時点における当該音片信号と前記音響信号との相互相関を、前記2以上の音片信号について累積した累積相互相関が最大化されるように、前記複数の音片信号の時系列を推定する。以上の態様では、各音片信号の末尾の時点における相互相関を複数の音片信号の時系列にわたり累積した累積相互相関が最大化されるように、複数の音片信号の時系列が推定される。したがって、各音片信号と音響信号との波形の類似性という観点から、音響信号を構成する複数の音片信号Qの時系列を高精度に推定できるという利点がある。
(6) From the form illustrated above, for example, the following configuration is grasped.
<Aspect 1>
In the acoustic analysis method according to a preferred aspect (aspect 1) of the present invention, each of the N sound piece signals is compared with the sound signal, so that the time series of the plurality of sound piece signals constituting the sound signal is obtained. presume. According to the above configuration, by comparing each of the N sound piece signals with the sound signal, it is possible to estimate a time series of a plurality of sound piece signals constituting the sound signal.
<Aspect 2>
In a preferred example of aspect 1 (aspect 2), the time series estimation of the plurality of sound piece signals is performed by calculating a cross-correlation between the sound piece signal and the acoustic signal for each of the N sound piece signals. A correlation analysis and an estimation process for estimating a time series of the plurality of sound piece signals selected from the N sound piece signals. In the estimation process, two or more sound pieces arranged on a time axis A time series of the plurality of sound piece signals so that a cumulative cross-correlation of the sound signal and the sound signal at the end of the signal with respect to the two or more sound piece signals is maximized. Is estimated. In the above aspect, the time series of the plurality of sound signal is estimated so that the accumulated cross correlation obtained by accumulating the cross correlation at the end of each sound signal over the time series of the plurality of sound signal is maximized. The Therefore, there is an advantage that the time series of the plurality of sound piece signals Q constituting the sound signal can be estimated with high accuracy from the viewpoint of the waveform similarity between each sound piece signal and the sound signal.

<態様3>
本発明の好適な態様(態様3)に係る音響解析装置は、N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する解析処理部を具備する。以上の構成によれば、N個の音片信号の各々と音響信号とを対比することで、音響信号を構成する複数の音片信号の時系列を推定することが可能である。
<Aspect 3>
The acoustic analysis device according to a preferred aspect (aspect 3) of the present invention compares each of the N sound piece signals with the sound signal, thereby obtaining a time series of a plurality of sound piece signals constituting the sound signal. An analysis processing unit for estimation is provided. According to the above configuration, by comparing each of the N sound piece signals with the sound signal, it is possible to estimate a time series of a plurality of sound piece signals constituting the sound signal.

100…音響解析装置、200…情報提供装置、300…端末装置、12,32…制御装置、14,34…記憶装置、20…解析処理部、22…相関解析部、24…推定処理部、36…収音装置、38…放音装置、42…変調処理部、44…混合処理部。
DESCRIPTION OF SYMBOLS 100 ... Acoustic analysis apparatus, 200 ... Information provision apparatus, 300 ... Terminal device, 12, 32 ... Control apparatus, 14, 34 ... Storage device, 20 ... Analysis processing part, 22 ... Correlation analysis part, 24 ... Estimation processing part, 36 ... sound collecting device, 38 ... sound emitting device, 42 ... modulation processing unit, 44 ... mixing processing unit.

Claims (3)

N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する
音響解析方法。
An acoustic analysis method for estimating a time series of a plurality of sound piece signals constituting the sound signal by comparing each of the N sound piece signals with the sound signal.
前記複数の音片信号の時系列の推定は、
前記N個の音片信号の各々について、当該音片信号と前記音響信号との相互相関を算定する相関解析と、
前記N個の音片信号から選択した前記複数の音片信号の時系列を推定する推定処理とを含み、
前記推定処理においては、時間軸上に配列された2以上の音片信号の末尾の時点における当該音片信号と前記音響信号との相互相関を、前記2以上の音片信号について累積した累積相互相関が最大化されるように、前記複数の音片信号の時系列を推定する
請求項1の音響解析方法。
Time series estimation of the plurality of sound piece signals is
For each of the N sound piece signals, a correlation analysis for calculating a cross-correlation between the sound piece signal and the acoustic signal;
An estimation process for estimating a time series of the plurality of sound piece signals selected from the N sound piece signals,
In the estimation process, a cross-correlation between the sound signal and the acoustic signal at the end of the two or more sound signal arranged on the time axis is accumulated for the two or more sound signals. The acoustic analysis method according to claim 1, wherein a time series of the plurality of sound piece signals is estimated so that the correlation is maximized.
N個の音片信号の各々と音響信号とを対比することで、前記音響信号を構成する複数の音片信号の時系列を推定する解析処理部
を具備する音響解析装置。
An acoustic analysis apparatus comprising: an analysis processing unit that estimates a time series of a plurality of sound piece signals constituting the sound signal by comparing each of the N sound piece signals with the sound signal.
JP2017044432A 2017-03-08 2017-03-08 Acoustic analysis method and acoustic analyzer Active JP6841095B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2017044432A JP6841095B2 (en) 2017-03-08 2017-03-08 Acoustic analysis method and acoustic analyzer

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2017044432A JP6841095B2 (en) 2017-03-08 2017-03-08 Acoustic analysis method and acoustic analyzer

Publications (2)

Publication Number Publication Date
JP2018146901A true JP2018146901A (en) 2018-09-20
JP6841095B2 JP6841095B2 (en) 2021-03-10

Family

ID=63591990

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2017044432A Active JP6841095B2 (en) 2017-03-08 2017-03-08 Acoustic analysis method and acoustic analyzer

Country Status (1)

Country Link
JP (1) JP6841095B2 (en)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61258297A (en) * 1985-05-11 1986-11-15 日本電気株式会社 Trouble diagnosing apparatus for voice synthesizer
JP2003255930A (en) * 2002-03-06 2003-09-10 Dainippon Printing Co Ltd Audio signal encoding method
WO2008062782A1 (en) * 2006-11-20 2008-05-29 Nec Corporation Speech estimation system, speech estimation method, and speech estimation program

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61258297A (en) * 1985-05-11 1986-11-15 日本電気株式会社 Trouble diagnosing apparatus for voice synthesizer
JP2003255930A (en) * 2002-03-06 2003-09-10 Dainippon Printing Co Ltd Audio signal encoding method
WO2008062782A1 (en) * 2006-11-20 2008-05-29 Nec Corporation Speech estimation system, speech estimation method, and speech estimation program

Also Published As

Publication number Publication date
JP6841095B2 (en) 2021-03-10

Similar Documents

Publication Publication Date Title
AU2015297648B2 (en) Terminal device, information providing system, information presentation method, and information providing method
JP6276453B2 (en) Information providing system, program, and information providing method
AU2015297647B2 (en) Information management system and information management method
EP3223274A1 (en) Information provision method and information provision device
CN109032870A (en) Method and apparatus for test equipment
CN110324726A (en) Model generation, method for processing video frequency, device, electronic equipment and storage medium
CN105161116A (en) Method and device for determining climax fragment of multimedia file
CN106840209A (en) Method and apparatus for testing navigation application
WO2018005202A1 (en) Audio augmented reality system
CN104143340B (en) A kind of audio frequency assessment method and device
KR20160056104A (en) Analyzing Device and Method for User&#39;s Voice Tone
CN110324657A (en) Model generation, method for processing video frequency, device, electronic equipment and storage medium
JP6841095B2 (en) Acoustic analysis method and acoustic analyzer
Catellier et al. Wenets: A convolutional framework for evaluating audio waveforms
Sobieraj et al. Orthogonality-regularized masked NMF for learning on weakly labeled audio data
EP3761623A1 (en) Information processing method, information processing device, and program
Siki et al. Time-frequency analysis on gong timor music using short-time fourier transform and continuous wavelet transform
US11362745B2 (en) Radio wave state analysis method
JP5611393B2 (en) Delay time measuring apparatus, delay time measuring method and program
JP6614395B2 (en) Information providing method and information providing apparatus
CN114333909A (en) Emotion recognition method and device, computer equipment and storage medium
US20230410821A1 (en) Sound processing method and device using dj transform
US11495200B2 (en) Real-time speech to singing conversion
Das et al. Music source separation: A guide
JP7159674B2 (en) Information processing device and information processing method

Legal Events

Date Code Title Description
A621 Written request for application examination

Free format text: JAPANESE INTERMEDIATE CODE: A621

Effective date: 20200124

A977 Report on retrieval

Free format text: JAPANESE INTERMEDIATE CODE: A971007

Effective date: 20201007

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20201013

A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20201116

TRDD Decision of grant or rejection written
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20210119

A61 First payment of annual fees (during grant procedure)

Free format text: JAPANESE INTERMEDIATE CODE: A61

Effective date: 20210201

R151 Written notification of patent or utility model registration

Ref document number: 6841095

Country of ref document: JP

Free format text: JAPANESE INTERMEDIATE CODE: R151

S531 Written request for registration of change of domicile

Free format text: JAPANESE INTERMEDIATE CODE: R313532

R350 Written notification of registration of transfer

Free format text: JAPANESE INTERMEDIATE CODE: R350