Please use this identifier to cite or link to this item:
https://er.chdtu.edu.ua/handle/ChSTU/9923| Title: | Analysis of methods and approaches to audio data processing |
| Other Titles: | Аналіз методів та підходів до задачі обробки аудіоданих |
| Authors: | Radin, Vladyslav Riabyi, Myroslav Радін, Владислав Рябий, Мирослав |
| Keywords: | speech-to-text;automatic speech recognition;Whisper;Wav2Vec;audio stream processing;word error rate;neural networks;автоматичне розпізнавання мовлення;обробка аудіопотоків;нейронні мережі |
| Issue Date: | 2026 |
| Publisher: | Вісник Черкаського державного технологічного університету |
| Abstract: | Within the research, it is necessary to assess the automatic speech recognition systems available today,
adapting them to the Ukrainian language and considering the trade-offs between accuracy, performance, and
various other factors. The aim of the research was to compare the existing methods and technologies for speech
recognition, with the goal of creating a system for assessing the informational impact of the analysed audio files.
The study utilised traditional hidden Markov processes and Gaussian mixtures models, hybrid neural network
models, such as DeepSpeech, Wav2Vec 2.0, and Whisper, cloud-based solutions, such as those by Google, Amazon,
Microsoft, and IBM. Audio processing was carried out using speech detection and mel-frequency coefficients;
word recognition accuracy, processing time on the CPU and GPU, response latency and instability, and the impact
of recognition errors on the natural language processing pipeline with tokenisation, topic classification and
sentiment analysis was assessed. As a result, it became possible to establish the dependence of the accuracy
and performance of Automatic Speech Recognition systems upon the architecture of the recognition systems
and the hardware upon which they are installed. Classic Hidden Markov Models were found to have the lowest
performance in recognising the words spoken within audio files, with word error rates between 18 and 30%, as well
as processing times of 51 to 72 seconds, indicating their limited suitability for Ukrainian language recognition.
End-to-end neural network models, however, had significantly better recognition performance, with DeepSpeech
models having an error rate of 22%, Wav2Vec models having an error rate of 12%, and whisper models having an
error rate of only 7%. Models based upon deep learning and transformers were also found to be robust to phonetic
variations of the Ukrainian language. Furthermore, the local end-to-end neural network models had significantly
better processing speeds than either cloud-based solutions or classic Hidden Markov Models, with whisper models
taking only around 10 seconds to process an audio file. Cloud-based recognition systems had similar accuracy to
local models (between 7 and 10% error rate), but required dependence upon the network, indicating potential
privacy issues for those using such systems. Thus, results of this research indicate that whisper and Wav2Vec
models are the best methods to utilise for audio content analysis and information effect detection systems. The
findings of this research can be applied to the development of automatic transcription systems, audio monitoring
systems, voice analytics services, or the development of information influence detection modules. Актуальність дослідження визначається потребою комплексної оцінки сучасних систем автоматичного розпізнавання мовлення в умовах української мовної специфіки з урахуванням компромісу між точністю, продуктивністю, апаратними вимогами та впливом якості транскрипції на подальші інформаційно-аналітичні процеси. Мета дослідження полягала в порівнянні наявних методів і технологій розпізнавання мовлення для побудови системи оцінки інформаційного впливу. У дослідженні використовувалися традиційні моделі прихованих марковських процесів та гауссові суміші,гібридні нейронні мережі, локальні нейромережеві моделі DeepSpeech, Wav2Vec 2.0 і Whisper, хмарні сервіси розпізнавання мовлення від Google, Amazon, Microsoft та IBM. Здійснювалася обробка аудіо з використанням виявлення мовлення та мел-частотних коефіцієнтів, фіксувалися точність розпізнавання слів, час обробки на процесорі та графічному процесорі, затримка та нестабільність відповіді, а також оцінювався вплив помилок розпізнавання на конвеєр обробки природної мови з токенізацією,тематичною класифікацією та аналізом настрою. У ході дослідження встановлено істотну залежність точності та продуктивності Automatic Speech Recognition від архітектури моделей і апаратної підтримки. Класичні Hidden Markov Models продемонстрували найнижчу ефективність із Word Error Rate 18-30 % та часом обробки 51-72 с, що підтвердило їх обмежену придатність для української мови в умовах шуму й двомовного мовлення. Локальні end-to-end нейромережеві моделі забезпечили принципово кращі показники: DeepSpeech досяг 22 % Word Error Rate, Wav2Vec – близько 12 %, а Whisper Large-v3 – приблизно 7 %, що відповідає зменшенню кількості помилок у 3-4 рази. Було показано, що багатомовні трансформерні архітектури характеризуються підвищеною стійкістю до фонетичних варіацій української мови та кодсвічингу. Оптимізовані локальні моделі забезпечили обробку майже в реальному часі (Whisper Large-v3 ≈10 с на файл), суттєво перевершивши класичні підходи за швидкодією. Хмарні сервіси продемонстрували співставну точність (7-10 % Word Error Rate), однак виявили залежність від мережевих затримок і обмеження щодо конфіденційності. Узагальнено, Whisper і Wav2Vec визначено як оптимальну основу для систем аналізу аудіоконтенту та виявлення інформаційного ефекту, оскільки вони забезпечують найкращий баланс між точністю, продуктивністю та контролем над даними. Отримані результати можуть бути застосовані у розробці систем автоматичної транскрипції, моніторингу та аналізу аудіопотоків, розробці сервісів голосової аналітики та модулів раннього виявлення інформаційних впливів. |
| URI: | https://er.chdtu.edu.ua/handle/ChSTU/9923 |
| ISSN: | 2306-4412 (print) 2708-6070 (online) |
| DOI: | https://doi.org/10.62660/bcstu/2.2026.35 |
| Volume: | 31 |
| Issue: | 2 |
| First Page: | 35 |
| End Page: | 47 |
| Appears in Collections: | том 31, №2/2026 |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 5.pdf | 437.43 kB | Adobe PDF | ![]() View/Open | |
| зміст.pdf | 118.63 kB | Adobe PDF | ![]() View/Open | |
| титул.pdf | 198.1 kB | Adobe PDF | ![]() View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.


