Search for documents in a domain through Google. The objective is to extract metadata

Last update: Dec 16, 2022

Related tags

Overview

MetaFinder - Metadata search through Google

   _____               __             ___________ .__               .___                   
  /     \     ____   _/  |_  _____    \_   _____/ |__|   ____     __| _/   ____   _______  
 /  \ /  \  _/ __ \  \   __\ \__  \    |    __)   |  |  /    \   / __ |  _/ __ \  \_  __ \ 
/    Y    \ \  ___/   |  |    / __ \_  |     \    |  | |   |  \ / /_/ |  \  ___/   |  | \/ 
\____|__  /  \___  >  |__|   (____  /  \___  /    |__| |___|  / \____ |   \___  >  |__|    
        \/       \/               \/       \/               \/       \/       \/          
        
|_ Author: @JosueEncinar
|_ Description: Search for documents in a domain through Google. The objective is to extract metadata
|_ Usage: python3 metafinder.py -d domain.com -l 100 -o /tmp

Installation:

> pip3 install metafinder

Upgrades are also available using:

> pip3 install metafinder --upgrade

Usage

CLI

metafinder -d domain.com -l 20 -o folder [-t 10] [-v]

Parameters:

d: Specifies the target domain.
l: Specify the maximum number of results to be searched.
o: Specify the path to save the report.
t: Optional. Used to configure the threads (4 by default).
v: Optional. It is used to display the results on the screen as well.

In Code

import metafinder.extractor as metadata_extractor

documents_limit = 5
domain = "target_domain"
data = metadata_extractor.extract_metadata_from_google_search(domain, documents_limit)
for k,v in data.items():
    print(f"{k}:")
    print(f"|_ URL: {v['url']}")
    for metadata,value in v['metadata'].items():
        print(f"|__ {metadata}: {value}")

document_name = "test.pdf"
try:
    metadata_file = metadata_extractor.extract_metadata_from_document(document_name)
    for k,v in metadata_file.items():
        print(f"{k}: {v}")
except FileNotFoundError:
    print("File not found")

Author

This project has been developed by:

Josué Encinar García -- @JosueEncinar

Contributors

Félix Brezo Fernández -- @febrezo

Disclaimer!

This Software has been developed for teaching purposes and for use with permission of a potential target. The author is not responsible for any illegitimate use.

Search for documents in a domain through Google. The objective is to extract metadata

Related tags

Overview

MetaFinder - Metadata search through Google

Installation:

Usage

CLI

In Code

Author

Contributors

Disclaimer!

Owner

Josué Encinar

ConvBERT-Prod

This repository collects together basic linguistic processing data for using dataset dumps from the Common Voice project

🏖 Easy training and deployment of seq2seq models.

NLPShala , the best IDE for all Natural language processing tasks.

Image2pcl - Enter the metaverse with 2D image to 3D projections

Translate U is capable of translating the text present in an image from one language to the other.

Pre-Training with Whole Word Masking for Chinese BERT

CorNet Correlation Networks for Extreme Multi-label Text Classification

Practical Machine Learning with Python

List of GSoC organisations with number of times they have been selected.

A list of NLP(Natural Language Processing) tutorials built on Tensorflow 2.0.

Smart discord chatbot integrated with Dialogflow to manage different classrooms and assist in teaching!

ElasticBERT: A pre-trained model with multi-exit transformer architecture.

Question answering app is used to answer for a user given question from user given text.

STonKGs is a Sophisticated Transformer that can be jointly trained on biomedical text and knowledge graphs

Switch spaces for knowledge graph embeddings

MEDIALpy: MEDIcal Abbreviations Lookup in Python

Translation to python of Chris Sims' optimization function

Library of deep learning models and datasets designed to make deep learning more accessible and accelerate ML research.

A programming language with logic of Python, and syntax of all languages.