Acoustic Analysis and Unsupervised Clustering of Brazilian Portuguese Accents

DE LIMA, Thales Aguiar and DA COSTA ABREU, Marjory (2026). Acoustic Analysis and Unsupervised Clustering of Brazilian Portuguese Accents. Acoustics, 8 (3): 60. [Article]

Documents
37914:1385768
[thumbnail of DeCostaAbreu-AcousticAnalysisAndUnsupervised(VoR).pdf]
Preview
PDF
DeCostaAbreu-AcousticAnalysisAndUnsupervised(VoR).pdf - Published Version
Available under License Creative Commons Attribution.

Download (1MB) | Preview
Abstract
Speech is a fundamental component of human communication and has become increasingly important with the widespread adoption of voice-based technologies, text messaging systems, Chatbots, and Large Language Models (LLMs). While AI systems are going through exponential breakthroughs, they often prioritize the dominant variant of a language. Using the TEDx Talks Brazilian Accents (TTBAcc) dataset, this research provides an acoustic analysis of two Brazilian Portuguese accents: Nordestino and Paulistano and an unsupervised accent clustering to find distinctive characteristics and investigate factors that can potentially support the monitoring of emerging and endangered accents. The acoustic analysis compares rhotics and vowels produced by speakers of both accents. The analysis of rhotics provided limited evidence of meaningful difference between accents. In contrast, the vowels exhibit significant differences in the fundamental frequency (f0), the first and second formant frequencies (F1 and F2) and intensity, suggesting that these acoustic features may be more effective to distinguish between the investigated accents. For accent clustering, f0, length, volume, and formant frequencies (F1 and F2) are used as inputs. Among the evaluated approaches, K-Means, with K=38 achieves the best performance, obtaining a Rand Score of 0.775 ± 0.028 and 0.524 ± 0.032 homogeneity, the best performing model. The optimal number of cluster exceeds the number of classified accents, which may indicate either an over-segmentation of data or a finer-grained phonetic variation captured in acoustic features. These findings highlight the potential of using acoustic features for investigating the variation in regional accents, and provides a basis for further research in monitoring emerging and endangered accents.
More Information
Statistics

Downloads

Downloads per month over past year

View more statistics

Metrics

Altmetric Badge

Dimensions Badge

Share
Add to AnyAdd to TwitterAdd to FacebookAdd to LinkedinAdd to PinterestAdd to Email

Actions (login required)

View Item View Item