Page to PAGE Layout Analysis Tool

Last update: Nov 24, 2022

Overview

P2PaLA

Page to PAGE Layout Analysis (P2PaLA) is a toolkit for Document Layout Analysis based on Neural Networks.

💥 Try our new DEMO for online baseline detection. ❗ ❗

If you find this toolkit useful in your research, please cite:

@misc{p2pala2017,
  author = {Lorenzo Quirós},
  title = {P2PaLA: Page to PAGE Layout Analysis tookit},
  year = {2017},
  publisher = {GitHub},
  note = {GitHub repository},
  howpublished = {\url{https://github.com/lquirosd/P2PaLA}},
}

Check this paper for more details Arxiv.

Requirements

Linux (OSX may work, but untested.).
Python (2.7, 3.6 under conda virtual environment is recomended)
Numpy
PyTorch (1.0). PyTorch 0.3.1 compatible on this branch
OpenCv (3.4.5.20).
NVIDIA GPU + CUDA CuDNN (CPU mode and CUDA without CuDNN works, but is not recomended for training).
tensorboard-pytorch (v0.9) [Optional]. pip install tensorboardX > A diferent conda env is recomended to keep tensorflow separated from PyTorch

Install

python setup.py install

To install python dependencies alone, use requirements file conda env create --file conda_requirements.yml

Usage

Input data must follow the folder structure data_tag/page, where images must be into the data_tag folder and xml files into page. For example:

mkdir -p data/{train,val,test,prod}/page;
tree data;

data
├── prod
│   ├── page
│   │   ├── prod_0.xml
│   │   └── prod_1.xml
│   ├── prod_0.jpg
│   └── prod_1.jpg
├── test
│   ├── page
│   │   ├── test_0.xml
│   │   └── test_1.xml
│   ├── test_0.jpg
│   └── test_1.jpg
├── train
│   ├── page
│   │   ├── train_0.xml
│   │   └── train_1.xml
│   ├── train_0.jpg
│   └── train_1.jpg
└── val
    ├── page
    │   ├── val_0.xml
    │   └── val_1.xml
    ├── val_0.jpg
    └── val_1.jpg

Run the tool.

python P2PaLA.py --config config.txt --tr_data ./data/train --te_data ./data/test --log_comment "_foo"

❗ Pre-trained models available here

Use TensorBoard to visualize train status:

tensorboard --logdir ./work/runs

xml-PAGE files must be at "./work/results/test/"

We recommend Transkribus or nw-page-editor to visualize and edit PAGE-xml files.

For detail about arguments and config file, see docs or python P2PaLa.py -h.
For more detailed example see egs:
- Bozen dataset see
- cBAD complex competition dataset see
- OHG dataset see

License

GNU General Public License v3.0 See LICENSE to see the full text.

Acknowledgments

Code is inspired by pix2pix and pytorch-CycleGAN-and-pix2pix

Page to PAGE Layout Analysis Tool

Related tags

Overview

P2PaLA

Requirements

Install

Usage

License

Acknowledgments

Owner

Lorenzo Quirós Díaz

Maze generator and solver with python

Fatigue Driving Detection Based on Dlib

Pixel art search engine for opengameart

Programa que viabiliza a OCR (Optical Character Reading - leitura óptica de caracteres) de um PDF.

An Implementation of the FOTS: Fast Oriented Text Spotting with a Unified Network

A bot that extract text from images using the Tesseract OCR.

textspotter - An End-to-End TextSpotter with Explicit Alignment and Attention

A Tensorflow model for text recognition (CNN + seq2seq with visual attention) available as a Python package and compatible with Google Cloud ML Engine.

This is a Computer vision package that makes its easy to run Image processing and AI functions. At the core it uses OpenCV and Mediapipe libraries.

Convolutional Recurrent Neural Networks(CRNN) for Scene Text Recognition

This tool will help you convert your text to handwriting xD

1st place solution for SIIM-FISABIO-RSNA COVID-19 Detection Challenge

This is a repository to learn and get more computer vision skills, make robotics projects integrating the computer vision as a perception tool and create a lot of awesome advanced controllers for the robots of the future.

Fun program to overlay a mask to yourself using a webcam

pyntcloud is a Python library for working with 3D point clouds.

Image processing is one of the most common term in computer vision

Code for the paper "Controllable Video Captioning with an Exemplar Sentence"

scantailor - Scan Tailor is an interactive post-processing tool for scanned pages.

Introduction to Augmented Reality (AR) with Python 3 and OpenCV 4.2.