Подобрена функция параметър извличане от реч сигнали използване на машина обучение алгоритъм (2)

May 30, 2023

3.7. Dataset Dataset collection and generation processes were performed as follow. In this study, the dataset with 120 h of audio was used for the model training. The dataset includes speech audio recordings which consist of sentences with a maximum of 15 words with a total length of approximately 120 h.

Desert living cistanche

Пустинен женшен

В допълнение, the dataset includes a large amount of different text for use in developing the language model. Over 90,650 utterances, 415,780 words, and 65,810 unique words that were included in the text corpus were collected, resulting in around 120 h of transcribed speech data. We split the dataset into training, validation, and test sets. The dataset statistics

са отчетени в Таблица 3.

Таблица 3. Наборът от данни спецификациите.

Table 3. The dataset specifications.

Ние разделяме набора от данни в три папки съответстващи на обучението 2c валидиране% 2c и тест комплекти. Всяка папка съдържа аудио записи и преписи. Аудио и съответстваща транскрипция имена на файлове са на същото% 2c с изключение на това, че аудио записи са съхранени като WAV файлове % 2c като има предвид транскрипциите са съхранени като TXT файлове използване на UTF% 7b% 7b0% 7d% 7d кодиране. Всички транскрипциите са представени използване на латиница азбука consisting на 29 букви и апостроф символ. Да предотврати пренапасване, ние приложени данни уголемяване техники базирани на скорост смущение и спектрални подобрения.

Cistanche deserticola slice (11)

Cistanche deserticola

4. Предложено Метод

An важно задача за програмисти когато разработване реч разпознаване системи е създаването на един оптимален метод за на параметър израз на реч сигнали % 5b45% 5d. Това метод активира отлично разделяне на звуци и изговорени думи докато осигуряване че говорители са нечувствителни към произношение модели и промени в акустична среда. Повечето грешки в дума разпознаване са причинени от а промяна в стъпката на сигнала дължими на а смяна в микрофона или а разликата в терена на произношението [46]. Друг чести причина на грешки е случайни нелинейни деформации на спектъра форма, които са винаги присъстват в речта сигнал на a говорител [47,48]. Следователно, one of the most important tasks in creating ефективно speech recognition systems is the selection of a representation that is sufficient for the content of the analyzed signal as well as insensitive to the voices of speakers and various acoustic environments.

Cistanche supplement near me—Improve memory2

Cistanche добавка близо ме-подобряване памет

The system that is used for extracting feature parameters typically has the following requirements. The information content, that is, the set of feature parameters, must ensure the reliable identification of recognizable speech elements. Furthermore, the loudness, that is, the maximum compression of the audio signal, and the non-statistical correlation of the parameters must be minimized. Independence from the speaker must also be achieved, that is, the maximal removal of information relating to the characteristics of the speaker from the vector of characters. Finally, homogeneity, which refers to the parameters having the same average variance and the ability to use simple metrics to determine the affinity between character sets, must be provided [49]. However, it is not always possible to satisfy all requirements simultaneously because such requirements are contradictory. The parametric description of the speech elements should be sufficiently detailed to distinguish them reliably and should be as laconic as possible.

Desert living cistanche

Супермен билки cistanche

В практика, the speech signal that is received from a microphone is digitized at a sampling rate of 8 to 22 kHz. Serial numerical values are divided into speech fragments (frames) with a duration of 10 to 30 ms, which correspond to quasi-stationary speech parts. A vector of features is computed from each frame, which is subsequently used at the acoustic level of speech recognition. At present, a wide range of methods is available for the parametric representation of signals based on autocorrelation analysis, hardware linear filtering, spectral analysis, and LPC. The most common approach for speech parameterization is the spectral analysis of signal fragments and the calculation of the the cepstral coefficients.

MFCCs have been used as informative features for the speech signal [41]. These characteristics are used extensively in speech recognition and are based on two main concepts: the cepstral and Mel scales. The main advantages of the algorithm are its high level of familiarity and ease of speech. The MFCC features are separated from the recorded speech signals. The MFCC algorithm uses the results of the phonogram and spectrum switching algorithms. The classical algorithm that is used to calculate the MFCCs is depicted in Figure 5.

Figure 5. Classical scheme for calculating MFCCs.

Figure 5. Classical scheme for calculating MFCCs.

Това проучване представя a бърз метод за извличане функцията параметри от a реч сигнал. На предложен алгоритъм за на бързо изчисление на на MFCCs е показано в Фигура 6.

Figure 6. Proposed framework for calculating MFCCs.

Фигура 6. Предложена рамка за изчисляване MFCCs.

Ние разгледа на изпълнение последователност на на предложен алгоритъм за на бързо extrac% 02tion на функция параметри от а реч сигнал

4.1. Разделение в Рамки

Следване предварително филтриране% 2c речта сигнал е разделено на 16 ms рамки. Всеки кадър (с изключение на за първия ) съдържа финала 10 ms на предишния кадър. Това процес продължава до края на сигнала В това проучване% 2c защото на вземане на проби скорост на на речта сигнал е 16 kHz% 2c на рамка дължина е N % 7b% 7b3% 7d% 7d% 2c и отместването дължина е M % 7b% 7b4% 7d% 7d.е 62.5 процента от кадъра дължина. A покритие от 50 процента до 75 процента от рамката дължина е общо препоръчително.

4.2. Hanning Прозорец и Намаляване Стойности

A Hanning прозорец размер на 1D беше използван. Ханинг прозорец е също наречен повдигнат косинус прозорец Ханинг прозорец може да бъде мисъл на като сумата на честотата спектър на три правоъгълна време прозорци. Това може да използва страната лобове да отмени всеки друг вън% 2c елиминиране висока честота смущения и енергия изтичане. Hanning прозорци са много полезен прозорец функции.% c2% a0

cistanche—Improve memory4

Cistanche добавка близо ме-подобряване памет

Кликни тук за изглед  Cistanche Подобряване Памет и Предотвратяване Алцхаймер% 27s Болест продукти

【Ask for more】 Email:cindy.xue@wecistanche.com /  Whats App:  {0086 18599088692 /  Wechat:  18599088692 

A тегло кутия е използва да намали изкривяване и гладка на индивидуално рамки. На плаващ сигнал че е разгледан в това проучване се състои на а плътен тон. На величина на на плосък тон е определя от на величина на чист тон в честота f % 2c което е филтрирано през на Hanning прозорец. 

A critical aspect of this window is that it sets the borders of the frames to zero. In this case, the short energies can be calculated while passing through the window, and they can be transferred from the sequence that is used to calculate the lowest amplitude energies. The aim is to remove low-energy signals from the signal by calculating the signal energy while smoothing the signal from this window. This process requires the following signal energy equation:


image

където En е на енергия на на вход сигнал фрагмент % 2c и xi е на сигнал стойност.

In addition to the window process using (11), the signal is processed in the next step, which signifificantly reduces the number of values that enter the processor. Figure 7 presents the parallel processing algorithm.

Прозорецът размерът представлява a брой на проби и a продължителност. Това е на основен параметър на на анализ. Прозорецът размер зависи на фундаменталната честота % 2c интензивност и промени в сигнала.

image

Figure 7. Hanning window algorithm for removing silent parts.
4.3. Short-Time Fourier Transform (STFT) Switches An intuitive understanding exists of the meaning of a high or low height. STFT
is a Fourierrelated transform that is used to determine the sinusoidal frequency and phase content of local sections of a signal as it changes over time. In practice, the STFT computation involves the division of a longer time signal into shorter segments of equal length, ollowed by a separate calculation of the Fourier transform on each shorter segment.This reveals the Fourier spectrum of each shorter segment.Discrete-time signals are used in practice. The corresponding time-frequency conver-sion is a discrete Fourier conversion, which describes the length of signal Xn as representative of the complex value frequency domain of the N coefficients. STFT, which describesthe evolution of the frequency components over time, is one of the most widely used toolsfor speech analysis and processing [50]. Similar to the spectrum itself, one advantage ofSTFT is that its parameters have physical and intuitive interpretations. STFT is typicallyvisualized using the log spectra 20log10 (X(h, k)). Such 2D log spectra can then be viewedusing a thermal map known as a spectrogram.
In the third stage of the algorithm, the STFT spectral switching procedure is applied tothe frames that are passed through the weight window. The STFT of the signal is obtainedby opening the windows and determining the DFT of each window. In particular, the
transformation for the Xn and Wn windows of the input signal is determined as follows:

image

където на k-индекс съответства на честотата стойности, и wn is прозореца функция, което е често a Hanning window или Gaussian window that is центрирано около нула.

4.4. Мел Трансформ

В на четвърти етап% 2c на сигнал че е прехвърлен към на честота лента е разделен в диапазони използване на триъгълни fifilters . На fifilter граници са изчислени използване на тебешир честота. На преход към на тебешир честота fifield е въз основа на на на следното уравнение :

image

където f е на честота диапазон.

На обратен превключвател е определено като следва:

image

Consider NN as the number of fifilters (26 fifilters are generally used) and flflows, as high as the frequency range under study. This range is transferred to the Mel scale and divided into NN evenly distributed intersecting ranges. The linear frequency-appropriate boundaries are determined within the fifield, whereas the weighting coeffificients that are obtained based on fifiltration are denoted by H. Subsequently, the fifilters are applied to the square modulus of the coeffificients obtained from the Fourier transform. The values that are obtained are logarithmic owing to the following expression:

image

В единствено число стойност разлагане алгоритъм е изпълнява в на fifinal етап на изчисляване на MFCCs.

5. Experimental Results

Ние внедрени и тествани на предложен метод в Visual Studio 2019 C плюс плюс на а PC с а 4.90 GHz CPU% 2c 32 GB на RAM% 2c и две Nvidia GeForce 2080Ti графични процесори , като показано в таблица 4. Системата е тестван в различни устройство среди да оцени производителността на сигнала функция извличане метод. In the experiments, during the operation of the algorithm (Figure 8), when the sequence of surfaces of the signal was n=15, the number of values was reduced by 40–50 percent , and the processing time increase 1.2-fold. Следователно, this algorithm exhibited signifificantly higher effificiency. Moreover, this algorithm enabled the separation and elimination of areas of silence while passing through прозореца Ханинг .

Several possible means are available for avoiding the waste of memory bandwidth. We propose a new solution that increases computing performance by matching the size of the signal frames to a block size of cache memory. This type of optimization can signifificantly affect the overall parallel processing performance. However, it can be used in digital signal processing by dividing the signal into frames through implementations on multicore processors. However, in practice, the selection is usually performed on a small scale, corresponding to the width of the data bus that connects the cache memory to the main memory and the size of its block. Our method implements the optimal use of these memories in parallel computing. The organization of the cache memory plays an essential role in parallel processing algorithms when dividing data into streams. In particular, the presence of vector-matrix effects in digital signal processing and the size of their streams should be adjusted according to the size of the cache blocks. This can be achieved using the proposed method, as illustrated in Figure 9.

Таблица 4. The подробни спецификации на на експериментални настройка.

image

image

фигура 8. (а) Първоначално входящо сигнал и (б) външен вид на сигнал след прилагане предложено алгоритъм.

image

Фигура 9. Паралел изчислителна структура използване на RK3288 процесори.

В това раздел%2в ние обсъждаме а количествено анализ до сравнявам на изпълнение на дифферент системи. Ние сравняваме нашите метод с добре познатите речта разпознаването алгоритми базирани на дълбокото обучение подходи. Оценка метрики са съществени за компутинг различни стратегии за реч разпознаване и оценяване изпълнението на различните подходи. Въпреки че ние използвахме резултатите от други проучвания за сравнение% 2c ние не сме сигурни дали те са верни защото източникът кодове и набори от данни на тези методи не са публично достъпни за проверка действителното изпълнение. Фигура 10 изобразява резултата от removing the silent части по време на преминаването a реч сигнал фрагмент през Хан% 02ning прозорец въз основа на на предложен бърз алгоритъм Скоростта резултати получени от анализа са представени в Таблица 5.

image

Фигура 10. k стойност на KNN алгоритъм (с функция селекция).

Table 5. Experimental results of proposed method.

image

A диапазон е е бил определен в ред да fifind степента на квартала това дава най-добрата точност стойност в KNN алгоритъм. В специфициран диапазон обхваща 1% e2% 80% 9325. In Фигура 10% 2c графиката на KNN алгоритъм с функция избор е приложен. Когато графиката е била разгледана , когато квартала стойност е била 1 в началото % 2c обучението точност е била много по-висока от теста точност. В на KNN алгоритъм създаден използване на функции избран от корелация% 2c на точността на модела е е определено да да бъде 99.15 процент в на обучение набор от данни и 97.35 процент в тест набор от данни.

The word error rate (WER) or character error rate is typically used to evaluate the accuracy of feature extraction from a speech signal. These are objective matrices that are helpful for a fair comparison of recognition techniques. In our previous studies [51–56], we computed metrics such as the F-measure (FM), precision, and recall. The FM is the weighted average that balances the measurements between the precision and recall проценти. Точността е съотношението на числото на правилно прогнозирано положително наблюдения към общото числото от прогнозираното положително наблюдения. Отзоваването е съотношението на числото на правилно прогнозирано положително наблюдения към общото число на наблюденията в действително клас, като посочено в (9). Следването уравнения може бъде използвано до изчисляване средното прецизност и изземване проценти на характеристика извличане методи:

image

където TP означава числото на вярно положително, FP обозначава числото на невярно положително, и FN означава числото на невярно отрицателно.

The FM е изчислено използване (10) % 2c като се има предвид и двете точност и изземване.

image

The average FM, recall, and precision of the proposed method was 98.4 percent . False detection occur in 1.6 percent of cases дължими на the unwanted noise of signals at the microphone. The range of the model's accuracy was between 0} and 1, and the metric estimation scores reach reach their best values at 1. An evaluation of our method and other recently published speech feature extraction methods is представени в Таблица 6. Същите брой на функции бяха използвани за а справедливо сравнение. А общо от 325 реч извадки от всяка група бяха анализирани от субекти с подобни характеристика фонове. Ефектите на различни кадър дължини според до на брой на фифилтър банки в на МФКК и различни рамка дължини в реда на на ЗПК бяха също разгледани за подобрена точност.

Таблица 6. Количествени точност резултати на реч характеристика екстракция.

image

Както споменато преди това% 2c на WER е най-много обща мярка за реч разпознаване% 02tion производителност. Това е изчислен от сравняване a справка транскрипция с на изход на на реч разпознаване. Базирано на това сравнение, it е възможно до изчисляване на числото на грешки, което обикновено принадлежат към три категории: (1) вмъквания когато a дума е не присъстват в препратката в изхода на на автоматичните реч разпознаване (ASR), (2) изтривания когато a дума е пропуснати в на ASR изход, и (3) замествания когато a дума е объркана с друга дума. The WER може бъде изчислена като следва.

image

където S е числото на заместванията на думи че са неправилно разпознати, D е числото на изтриванията, I е числото на вмъкванията, и N е числото на думите в препратката транскрипцията. Основният въпросът в изчислителните този резултатът е подравняването между на последователностите Това може бъде определено чрез динамично програмиране използвайки на Levenshtein разстояние [67].

Въз основа на Таблица 6% 2c ние извършени а статистически анализ да посочи средната стойност на на сравнени методи използване на WER оценка метрика% 2c като илюстрирани в Фигура 11. На подобрена функция екстрактор даде един точност на приблизително 98.4 процент % 2c като има предвид на други подходи даде точност между 78 процент и 96 процент . Ние използвахме резултатите предоставихме в съответните документи за сравнение% 3b обаче% 2c точността на тези стойности не е лесно проверими защото източникът кодове и набори от данни на тези методи не са публично достъпни до потвърди тяхната реална производителност. Въпреки това, in the case of standard scenes, the proposed method was experimentally demonstrated to provide excellent speech feature extraction accuracy by reducing the computational time, even when the speech data are noisy or of low quality.

Figure 11. Quantitative results of speech signal feature extraction approaches using vertical graphs.

Фигура 11. Количествени резултати на реч сигнал функция екстракция подходи използване на вертикални графики.

Освен това, we evaluated the false-positive results of the selected methods. As can be observed from Figure 12, the proposed approach had the fewest errors. Furthermore, the highly effificient parallel computation method signifificantly reduced the sound signal feature selection and extraction errors. Overfifitting was one of the main issues during the training, and almost all machine learning models suffer from it. Ние се опитахме да намалим overfifit риск използване на a функция подбор техника че цели вместо да ранг на важност на съществуващите функции в на набор от данни и изхвърли по-малко важни такива (не нови функции са създадени).

image

Фигура 12. Видими резултати на фалшиво положителни реч сигнал функция екстракция експерименти.

Таблица 7 показва изпълнението резултатите на методите които бяха използвани в реч разпознаване среди базирани на различни свойства. Нашите предложени подход не не страдат от нежелани и ненужни фон шум и е не засегнати от нискокачествениity човешки гласове такива като дрезгави гласове, гласове произведени с a възпалено гърло, или дори звуци от хора с пълна глас загуба. Нашите метод цели да преодолеят на вероятносттаlems of неадекватен запис оборудване, фон шум, трудно акценти и диалекти, и различни височини в a глас. В а нормална среда, най-добрите резултати за точно откриване и извличане реч характеристика предизвикателства бяха получени използващи на предложени метод с а намалена обработка време.

Table 7. Review of speech feature detection and extraction performance using various features.

image

Резултатите от речта разпознаване методи бяха класифицирани като мощни% 2c нормални % 2c или слаби за седемте категории. Мощният критерий демонстрира че алгоритъмът може да преодолее всички типове от предизвикателства. В контраст % 2c нормалния критерий показва, че алгоритъмът може да се провали в определени случаи защото думата границите не са дефинирани предварително. И накрая% 2c слабия критерий предполага, че алгоритъмът е ненадежден под фон шум или вибрации.

6. Ограничения

It is diffificult to conclude that the methods proposed to date do not exhibit any shortcomings. Our proposed method may also result in errors дължими to various noise environments. To overcome this problem, we aimed to reduce the number of features in the dataset by create new features from existing features [69]. Тъй като overfifitting беше един от основните въпроси за обучение различни модели по време на конкуренцията % 2c обогатяване на обучение данни от% c2% a0adding данни проби от различни ресурси може да бъде a възможно решение за подобряване резултатите . Irrespective на на гореспоменатите проблеми% 2c експерименталните резултати разкри че наш метод е много здрав и ефективен за реч функция извличане задачи , с средно точност на 98.4 процента и FM от 99.5 процента .

7. Conclusions

A нова висока производителност паралелно изчислителна техника подход използване на а машина обучение метод е бил предложен за реч разпознаване системи. Ускорение проблеми в машини с ограничени изчислителни ресурси може да бъде решен използване на разпределени системи. Изчисление скорост в сигнал разпознаване системи може да бъде увеличен , и производителност на многоядрени платформи може да бъде подобрена от създаване и използване на ефикасност и бързо algorithms. Резултатите демонстрират че предложеният модел намалява обработката времето и подобрява функцията извличането точността с 98.4 процента ефективно използване на MFCCs . Това е било наблюдавано че характеристиките с ниска корелация стойности извлечени от функция избор са също ефективни в успеха на модела . Статистическият анализ беше извършен на предварително обработените данни , и смислена информация беше произведена от данните използване на данни използване на K-най-близо съседи (KNN) машина обучение алгоритъм.

Future studies will focus on improving the accuracy of our method by using deep learning approaches and optimizing the cache memory of multicore processors to detect and extract speech signals without a signifificant quality loss. Furthermore, we plan to construct a spectral analysis model based on parallel processing with robust analysis performance that will enable the establishment of embedded devices with low computational resources using Taris speech datasets [70] in the 3D CNN and 3D U-Net Environment [71–75]

1. Meng, Y.J.; Liu, W.J.; Zhang, R.Z.; Du, H.S. Speech Feature Parameter Extraction and Recognition Based on Interpolation. Appl. Mech. Mater. 2014, 602–605, 2118–2123. [CrossRef] 2. Musaev, M.; Rakhimov, M. Accelerated Training for Convolutional Neural Networks. In Proceedings of the 2020 International Conference on Information Science and Communications Technologies (ICISCT), Tashkent, Uzbekistan, 4–6 November 2020; pp. 1–5. [CrossRef] 3. Ye, F.; Yang, J. A Deep Neural Network Model for Speaker Identifification. Appl. Sci. 2021, 11, 3603. [CrossRef] 4. Musaev, M.; Rakhimov, M. A Method of Mapping a Block of Main Memory to Cache in Parallel Processing of the Speech Signal. In Proceedings of the 2019 International Conference on Information Science and Communications Technologies (ICISCT), Karachi, Pakistan, 9–10 March 2019; pp. 1–4. [CrossRef] 5. Jiang, N.; Liu, T. An improved speech segmentation and clustering algorithm based on SOM and k-means. Math. Probl. Eng. 2020, 2020, 3608286. [CrossRef] 6. Hu, W.; Yang, Z.; Chen, C.; Sun, B.; Xie, Q. A vibration segmentation approach for the multi-action system of numerical control turret. Signal Image Video Process. 2021, 16, 489–496. [CrossRef]

7. Popescu, T.D.; Aiordachioaie, D. Fault detection of rolling element bearings using optimal segmentation of vibrating signals. Mech. Syst. Signal Process. 2019, 116, 370–391. [CrossRef] 8. Shihab, M.S.H.; Aditya, S.; Setu, J.H.; Imtiaz-Ud-Din, K.M.; Efat, M.I.A. A Hybrid GRU-CNN Feature Extraction Technique for Speaker Identifification. In Proceedings of the 2020 23rd International Conference on Computer and Information Technology (ICCIT), Dhaka, Bangladesh, 19–21 December 2020; pp. 1–6. [CrossRef] 9. Korkmaz, O.; Atasoy, A. Emotion recognition from speech signal using mel-frequency cepstral coefficients. In Proceedings of the 9th International Conference on Electrical and Electronics Engineering (ELECO), Bursa, Turkey, 26–28 November 2015; pp. 1254–1257. 10. Ayvaz, U.; Gürüler, H.; Khan, F.; Ahmed, N.; Whangbo, T.; Abdusalomov, A. Automatic Speaker Recognition Using Mel-Frequency Cepstral Coeffificients Through Machine Learning. CMC-Comput. Mater. Contin. 2022, 71, 5511–5521. 11. Al-Qaderi, M.; Lahamer, E.; Rad, A. A Two-Level Speaker Identifification System via Fusion of Heterogeneous Classififiers and Complementary Feature Cooperation. Sensors 2021, 21, 5097. [CrossRef] 12. Batur Dinler, Ö.; Aydin, N. An Optimal Feature Parameter Set Based on Gated Recurrent Unit Recurrent Neural Networks for Speech Segment Detection. Appl. Sci. 2020, 10, 1273. [CrossRef] 13. Kim, H.; Shin, J.W. Dual-Mic Speech Enhancement Based on TF-GSC with Leakage Suppression and Signal Recovery. Appl. Sci. 2021, 11, 2816. [CrossRef] 14. Lee, S.-J.; Kwon, H.-Y. A Preprocessing Strategy for Denoising of Speech Data Based on Speech Segment Detection. Appl. Sci. 2020, 10, 7385. [CrossRef] 15. Rusnac, A.-L.; Grigore, O. CNN Architectures and Feature Extraction Methods for EEG Imaginary Speech Recognition. Sensors 2022, 22, 4679. [CrossRef] [PubMed] 16. Wafa, R.; Khan, M.Q.; Malik, F.; Abdusalomov, A.B.; Cho, Y.I.; Odarchenko, R. The Impact of Agile Methodology on Project Success, with a Moderating Role of Person's Job Fit in the IT Industry of Pakistan. Appl. Sci. 2022, 12, 10698. [CrossRef] 17. Aggarwal, A.; Srivastava, A.; Agarwal, A.; Chahal, N.; Singh, D.; Alnuaim, A.A.; Alhadlaq, A.; Lee, H.-N. Two-Way Feature Extraction for Speech Emotion Recognition Using Deep Learning. Sensors 2022, 22, 2378. [CrossRef] [PubMed] 18. Marini, M.; Vanello, N.; Fanucci, L. Optimising Speaker-Dependent Feature Extraction Parameters to Improve Automatic Speech Recognition Performance for People with Dysarthria. Sensors 2021, 21, 6460. [CrossRef] 19. Tiwari, S.; Jain, A.; Sharma, A.K.; Almustafa, K.M. Phonocardiogram Signal Based Multi-Class Cardiac Diagnostic Decision Support Syste. IEEE Access 2021, 9, 110710–110722. [CrossRef] 20. Mohtaj, S.; Schmitt, V.; Möller, S. A Feature Extraction based Model for Hate Speech Identifification. arXiv 2022, arXiv: 2201.04227. 21. Kuldoshbay, A.; Abdusalomov, A.; Mukhiddinov, M.; Baratov, N.; Makhmudov, F.; Cho, Y.I. An improvement for the automatic classifification method for ultrasound images used on CNN. Int. J. Wavelets Multiresolution Inf. Process. 2022, 20, 2150054. 22. Passricha, V.; Aggarwal, R.K. A hybrid of deep CNN and bidirectional LSTM for automatic speech recognition. J. Intell. Syst. 2020, 29, 1261–1274. [CrossRef] 23. Mukhamadiyev, A.; Khujayarov, I.; Djuraev, O.; Cho, J. Automatic Speech Recognition Method Based on Deep Learning Approaches for Uzbek Language. Sensors 2022, 22, 3683. [CrossRef] [PubMed] 24. Li, F.; Liu, M.; Zhao, Y.; Kong, L.; Dong, L.; Liu, X.; Hui, M. Feature extraction and classifification of heart sound using 1D convolutional neural networks. EURASIP J. Adv. Signal Process. 2019, 2019, 59. [CrossRef] 25. Chang, L.-C.; Hung, J.-W. A Preliminary Study of Robust Speech Feature Extraction Based on Maximizing the Probability of States in Deep Acoustic Models. Appl. Syst. Innov. 2022, 5, 71. [CrossRef] 26. Ramírez, J.; Górriz, J.M.; Segura, J.C. Voice Activity Detection. Fundamentals and Speech Recognition System Robustness. In Robust Speech Recognition and Understanding; Grimm, M., Kroschel, K., Eds.; I-TECH Education and Publishing: London, UK, 2007; pp. 1–22. 27. Oh, S. DNN Based Robust Speech Feature Extraction and Signal Noise Removal Method Using Improved Average Prediction LMS Filter for Speech Recognition. J. Converg. Inf. Technol. 2021, 11, 1–6. [CrossRef] 28. Abbaschian, B.J.; Sierra-Sosa, D.; Elmaghraby, A. Deep Learning Techniques for Speech Emotion Recognition, from Databases to Models. Sensors 2021, 21, 1249. [CrossRef] 29. Rakhimov, M.; Mamadjanov, D.; Mukhiddinov, A. A High-Performance Parallel Approach to Image Processing in Distributed Computing. In Proceedings of the 2020 IEEE 14th International Conference on Application of Information and Communication Technologies (AICT), Uzbekistan, Tashkent, 7–9 October 2020; pp. 1–5. [CrossRef] 30. Abdusalomov, A.; Mukhiddinov, M.; Djuraev, O.; Khamdamov, U.; Whangbo, T.K. Automatic Salient Object Extraction Based on Locally Adaptive Thresholding to Generate Tactile Graphics. Appl. Sci. 2020, 10, 3350. [CrossRef] 31. Abdusalomov, A.; Whangbo, T.K. An improvement for the foreground recognition method using shadow removal technique for indoor environments. Int. J. Wavelets Multiresolution Inf. Process. 2017, 15, 1750039. [CrossRef] 32. Abdusalomov, A.; Whangbo, T.K. Detection and Removal of Moving Object Shadows Using Geometry and Color Information for Indoor Video Streams. Appl. Sci. 2019, 9, 5165. [CrossRef] 33. Mery, D. Computer Vision for X-ray Testing; Springer International Publishing: Cham, Switzerland, 2015; p. 271, ISBN 978-3319207469. 34. Mark, S. Speech imagery recalibrates speech-perception boundaries. Atten. Percept. Psychophys. 2016, 78, 1496–1511. [CrossRef] 35. Mudgal, E.; Mukuntharaj, S.; Modak, M.U.; Rao, Y.S. Template Based Real-Time Speech Recognition Using Digital Filters on DSP-TMS320F28335. In Proceedings of the 2018 Fourth International Conference on Computing Communication Control and Automation (ICCUBEA), Pune, India, 16–18 August 2018; pp. 1–6. [CrossRef]

Може да харесаш също