Skoči na glavni sadržaj

Izvorni znanstveni članak

https://doi.org/https://doi.org/10.22210/suvlin.2026.101.05

Croatian Language in the Transition from Neural Machine Translation to Large Language Models

Antoni Oliver orcid id orcid.org/0000-0001-8399-3770 ; Universitat Oberta de Catalunya *
Sergi Álvarez–Vidal ; Universitat Autònoma de Barcelona *

* Dopisni autor.


Puni tekst: engleski pdf 160 Kb

str. 101-120

preuzimanja: 0

citiraj


Sažetak

Machine translation (MT) technologies are currently undergoing a paradigm shift, transitioning from specialized Neural Machine Translation (NMT) frameworks to the broader capabilities of Large Language Models (LLMs). This paper examines the current standing of the Croatian language within this technological evolution.
While bilingual NMT models often exhibit high precision, multilingual NMT leverage transfer learning to enhance performance for low–resource language pairs, but with lower performance for high–resource ones. Conversely, LLMs—whether general–purpose or fine–tuned for translation—offer superior multilingual proficiency and context awareness. Unlike NMT, LLMs can process extended discourse, such as full paragraphs or documents, leading to significant improvements in coreference resolution and gender agreement. Despite the substantial computational requirements of LLMs, recent optimization techniques allow for smaller, more efficient versions that maintain high output quality.
This study evaluates the performance of various NMT and LLM architectures specifically for Croatian from/to English and Spanish using several automatic quality evaluation metrics. The findings demonstrate that open–source models can achieve, and occasionally surpass, the quality of Google Translate, a widely used commercial NMT system. Furthermore, while our evaluation focuses on this specific language triad, the multilingual nature of the analysed systems suggests that open–source models provide high–quality translation capabilities for Croatian across dozens, if not hundreds, of language pairs.

Ključne riječi

neural machine translation; large language models; Croatian; English; Spanish

Hrčak ID:

349604

URI

https://hrcak.srce.hr/349604

Datum izdavanja:

22.7.2026.

Podaci na drugim jezicima: hrvatski

Posjeta: 0 *