25 years of Hašek

Dembitz, Šandor

Jezik : časopis za kulturu hrvatskoga književnog jezika, Vol. 66 No. 4-5, 2019.

Izvorni znanstveni članak

25 years of Hašek

Šandor Dembitz

Puni tekst: hrvatski pdf 829 Kb

str. 138-150

preuzimanja: 278

citiraj

APA 6th Edition

Dembitz, Š. (2019). 25 years of Hašek. Jezik, 66 (4-5), 138-150. Preuzeto s https://hrcak.srce.hr/237727

MLA 8th Edition

Dembitz, Šandor. "25 years of Hašek." Jezik, vol. 66, br. 4-5, 2019, str. 138-150. https://hrcak.srce.hr/237727. Citirano 26.04.2024.

Chicago 17th Edition

Dembitz, Šandor. "25 years of Hašek." Jezik 66, br. 4-5 (2019): 138-150. https://hrcak.srce.hr/237727

Harvard

Dembitz, Š. (2019). '25 years of Hašek', Jezik, 66(4-5), str. 138-150. Preuzeto s: https://hrcak.srce.hr/237727 (Datum pristupa: 26.04.2024.)

Vancouver

Dembitz Š. 25 years of Hašek. Jezik [Internet]. 2019 [pristupljeno 26.04.2024.];66(4-5):138-150. Dostupno na: https://hrcak.srce.hr/237727

IEEE

Š. Dembitz, "25 years of Hašek", Jezik, vol.66, br. 4-5, str. 138-150, 2019. [Online]. Dostupno na: https://hrcak.srce.hr/237727. [Citirano: 26.04.2024.]

Sažetak

Hašek is a Croatian on-line spellchecker that continuously operates since March 21, 1994,
nowadays at the address https://ispravi.me/. In 25 years of functioning Hašek processed
nearly 30 million texts, which build a corpus of more than 7 billion tokens. By comparison,
all books ever published in Croatian form a corpus with less than 20 billion tokens.
As a WWW-embedded tool, Hašek took advantage of many web-based services including
learning. Thanks to Hašek’s learning capability, its dictionary increased from initial 100
thousand to more than 2 million word-types. Another aspect of learning was the creating
and regular updating of the Croatian n-gram system. Unlike Google, whose n-gram systems
are based on the WaC (Web as Corpus) approach and cut-off criteria, Croatian n-grams
were extracted from processed texts by a lexical criterion: each n-gram constituent must
be proven by the spellchecker as valid in Croatian spelling. The difference in approaches
made Croatian n-gram system comparable in size to the largest Google n-gram systems.
Unfortunately, the advantages of on-line spellchecking for rapid breakthroughs into much
more sophisticated language technology areas were not recognized by Croatian decision
makers, with some consequences mentioned in the paper.

Ključne riječi

Hašek; spellchecking; learning; Google; n-gram systems

Hrčak ID:

237727

URI

https://hrcak.srce.hr/237727

Datum izdavanja:

1.12.2019.

Podaci na drugim jezicima: hrvatski

Posjeta: 1.870 *

Prijava i registracija

Jezik : časopis za kulturu hrvatskoga književnog jezika, Vol. 66 No. 4-5, 2019.

Sažetak

Ključne riječi

Hrčak ID:

URI

Datum izdavanja: