Repository containing the code for An-Gocair text normaliser

Last update: Jun 28, 2022

Related tags

Overview

Scottish Gaelic Text Normaliser

The following project contains the code and resources for the Scottish Gaelic text normalisation project. The repo can be cloned and top level functions will allow you to normalise phrases or whole documents.

Installation

To use the program you will have to clone the repo and install dependencies in a python virtualenvironment using python 3 and above.

instructions

from GaelicTextNormaliser import TextNormaliser

normaliser = TextNormaliser(from_config="config.yaml")

normaliser.normalise_doc(doc="Bha rìgh òg Easaidh Ruagh an dèigh dha'n oighreachd fhaotainn da fèin ri mòran àbhachd, ag amharc a mach dè a chordadh ris,'s dè thigeadh r'a nadur.")

"Bha rìgh òg Easaidh Ruadh an dèidh dhan oighreachd fhaotainn da fhèin ri mòran àbhachd, ag amharc a-mach dè a chòrdadh ris,'s dè thigeadh ra nàdar."

Alternatively there is a webapp that can be found at https://www.garg.ed.ac.uk/an_gocair.

Acknowledgements

Scottish Gaelic Lexicon

The lexicon file is provided by Michael Bauer, Scottish Gaelic linguist, author and lead collaborator on the Am Faclaer Baeg SG dictionary. The lexicon is a reformatted version of the dictionary that makes use of Michael's extensive labelling of traditional Gaelic spellings and common misspellings. The resource is extremely vital for the success of the memory based approach.

Rules for Normalisation

The lexical and grammatical rules for normalisation were the result of collaboration between the project leader Dr Will Lamb and Baeur. Both Lamb and Bauer, as fluent Gaelic speakers and experienced proof readers, were able to provide the linguistic rules to be translated into python code.

Scottish Gaelic Part of Speech Tagger

For further conditioning in the rule based approach, part of speech tags were necessary. The code and models for POS tagging is very kindly provided by Loïc Boizou. The scripts were altered slightly to work within the python object.

Further Acknowledgements

This program was funded by the Data-Driven Innovation initiative (DDI), delivered by the University of Edinburgh and Heriot-Watt University for the Edinburgh and South East Scotland City Region Deal. DDI is an innovation network helping organisations tackle challenges for industry and society by doing data right to support Edinburgh in its ambition to become the data capital of Europe. The project was delivered by the Edinburgh Futures Institute (EFI), one of five DDI innovation hubs which collaborates with industry, government and communities to build a challenge-led and data-rich portfolio of activity that has an enduring impact.

Repository containing the code for An-Gocair text normaliser

Related tags

Overview

Scottish Gaelic Text Normaliser

Installation

instructions

Acknowledgements

Scottish Gaelic Lexicon

Rules for Normalisation

Scottish Gaelic Part of Speech Tagger

Further Acknowledgements

Owner

The app gets your sutitle.srt and proccess it to extract sentences

Fuzzy string matching like a boss. It uses Levenshtein Distance to calculate the differences between sequences in a simple-to-use package.

Search for terms(word / table / field name or any) under Snowflake schema names

box is a text-based visual programming language inspired by Unreal Engine Blueprint function graphs.

Build a translation program similar to Google Translate with Python programming language and QT library

一款高性能敏感词(非法词/脏字)检测过滤组件，附带繁体简体互换，支持全角半角互换，汉字转拼音，模糊搜索等功能。

This project aims to test check if your RegExp are being matched by grep.

A neat little program to read the text from the "All Ten Fingers" program, and write them back.

utoken is a multilingual tokenizer that divides text into words, punctuation and special tokens such as numbers, URLs, XML tags, email-addresses and hashtags.

text-to-speach bot - You really do NOT have time for read a newsletter? Now you can listen to it

This project is a small tool for processing url-containing texts delivered by HUAWEI Share on Windows.

StealBit1.1 and earlier strings and config extraction scripts

A python Tk GUI that creates, writes text and attaches images into a custom spreadsheet file

Word and phrase lists in CSV

An anthology of a variety of tools for the Persian language in Python

Export solved codewars kata challenges to a text file.

a python package that lets you add custom colors and text formatting to your scripts in a very easy way!

Python character encoding detector

Aml - anti-money laundering

A collection of pre-commit hooks for handling text files.