Vector space based Information Retrieval System for Text Processing - Information retrieval

Last update: Jan 01, 2022

Related tags

Text Processing BITS-IR-PROJECT

Overview

Information Retrieval: Text Processing

Group 13

Sequence of operations

Install Requirements
Add given wikipedia files to the corpus directory.
Download glove.6B.100d.txt dataset (Ignore if already present) and place it in the project root directory.
Run construct_index.py
Run construct_index.py --zoned_index True
Run trim_embeddings.py
Run test_queries.py
Run test_queries.py --score_title True
Run test_queries.py --expand_query True

Installing Requirements:

   pip install -r requirements.txt

corpus

Contains the files to be indexed. Add files directly to this directory. Do not create subdirectories.
For this assignment, we have used the following files present in the AA folder of wikipedia files.
wiki_00
wiki_01
wiki_05
wiki_06
wiki_10
wiki_11
wiki_15
wiki_16
wiki_20
wiki_21
wiki_25
wiki_26
wiki_30
wiki_31

index_files

Contains the inverted indices constructed using construct_index.py.

construct_index.py

Constructs the inverted indices used for query evaluation.
Command Line Arguments:
--zoned_index: True if zoned indexing must be used. Set to False by default.

trim_embeddings.py

Trims the GloVe embeddings to contain terms only from corpus. Download the glove.6B.100d.txt dataset before running this file.

test_queries.py

Evaluates queries and displays retrieved documents.
Command Line Arguments:
--score_title: True if zoned index considered for evaluation. Set to False by default.
--expand_query: True if query expansion must be used. Set to False by default.

helper_module.py

Contains helper functions used by other files. Do not run this file.

document_list.txt

Contains the document ids and names used for evaluation.

Vector space based Information Retrieval System for Text Processing - Information retrieval

Related tags

Overview

Information Retrieval: Text Processing

Group 13

Sequence of operations

Installing Requirements:

corpus

index_files

construct_index.py

trim_embeddings.py

test_queries.py

helper_module.py

document_list.txt

Owner

Extract price amount and currency symbol from a raw text string

从flomo导出的笔记中生成词云

A working (ish) python script to convert text to a gradient.

Chilean Digital Vaccination Pass Parser (CDVPP) parses digital vaccination passes from PDF files

Add your new words to a text file and get them randomly.

Open-source linguistic ethnography tool for framing public opinion in mediatized groups.

Free & simple way to encipher text

Skype export archive to text converter for python

The Levenshtein Python C extension module contains functions for fast computation of Levenshtein distance and string similarity

Answer some questions and get your brawler csvs ready!

You can encode and decode base85, ascii85, base64, base32, and base16 with this tool.

Correcting typos in a word based on the frequency dictionary

This is REST-API for Indonesian Text Summarization using Non-Negative Matrix Factorization for the algorithm to summarize documents and FastAPI for the framework.

Etranslate is a free and unlimited python library for transiting your texts

Migrates translations to the REDCap native Multi-Language Management system

Simple python program to auto credit your code, text, book, whatever!

A Python app which can convert normal text to Handwritten text.

Fuzz a language by mixing up only few words.

WorldCloud Orçamento de Estado 2022

PyNews 📰 Simple newsletter made with python 🐍🗞️