Geometric Augmentation for Text Image

Last update: Jan 05, 2023

Overview

Text Image Augmentation

A general geometric augmentation tool for text images in the CVPR 2020 paper "Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition". We provide the tool to avoid overfitting and gain robustness of text recognizers.

Note that this is a general toolkit. Please customize for your specific task. If the repo benefits your work, please cite the papers.

News

2020-02 The paper "Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition" was accepted to CVPR 2020. It is a preliminary attempt for smart augmentation.
2019-11 The paper "Decoupled Attention Network for Text Recognition" (Paper Code) was accepted to AAAI 2020. This augmentation tool was used in the experiments of handwritten text recognition.
2019-04 We applied this tool in the ReCTS competition of ICDAR 2019. Our ensemble model won the championship.
2019-01 The similarity transformation was specifically customized for geomeric augmentation of text images.

Requirements

GCC 4.8.*
Python 2.7.*
Boost 1.67
OpenCV 2.4.*

We recommend Anaconda to manage the version of your dependencies. For example:

     conda install boost=1.67.0

Installation

Build library:

    mkdir build
    cd build
    cmake -D CUDA_USE_STATIC_CUDA_RUNTIME=OFF ..
    make

Copy the Augment.so to the target folder and follow demo.py to use the tool.

    cp Augment.so ..
    cd ..
    python demo.py

Demo

Distortion

Stretch

Perspective

Speed

To transform an image with size (H:64, W:200), it takes less than 3ms using a 2.0GHz CPU. It is possible to accelerate the process by calling multi-process batch samplers in an on-the-fly manner, such as setting "num_workers" in PyTorch.

Improvement for Recognition

We compare the accuracies of CRNN trained using only the corresponding small training set.

Dataset	IIIT5K	IC13	IC15
Without Data Augmentation	40.8%	6.8%	8.7%
With Data Augmentation	53.4%	9.6%	24.9%

Citation

@inproceedings{luo2020learn,
  author = {Canjie Luo and Yuanzhi Zhu and Lianwen Jin and Yongpan Wang},
  title = {Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition},
  booktitle = {CVPR},
  year = {2020}
}

@inproceedings{wang2020decoupled,
  author = {Tianwei Wang and Yuanzhi Zhu and Lianwen Jin and Canjie Luo and Xiaoxue Chen and Yaqiang Wu and Qianying Wang and Mingxiang Cai}, 
  title = {Decoupled attention network for text recognition}, 
  booktitle ={AAAI}, 
  year = {2020}
}

@article{schaefer2006image,
  title={Image deformation using moving least squares},
  author={Schaefer, Scott and McPhail, Travis and Warren, Joe},
  journal={ACM Transactions on Graphics (TOG)},
  volume={25},
  number={3},
  pages={533--540},
  year={2006},
  publisher={ACM New York, NY, USA}
}

Acknowledgment

Thanks for the contribution of the following developers.

@keeofkoo

@cxcxcxcx

@Yati Sagade

Attention

The tool is only free for academic research purposes.

Geometric Augmentation for Text Image

Related tags

Overview

Text Image Augmentation

News

Requirements

Installation

Demo

Speed

Improvement for Recognition

Citation

Acknowledgment

Attention

Owner

Canjie Luo

一键翻译各类图片内文字

This is used to convert a string to an Image with Handwritten Characters.

Reference Code for AAAI-20 paper "Multi-Stage Self-Supervised Learning for Graph Convolutional Networks on Graphs with Few Labels"

This is the official PyTorch implementation of the paper "TransFG: A Transformer Architecture for Fine-grained Recognition" (Ju He, Jie-Neng Chen, Shuai Liu, Adam Kortylewski, Cheng Yang, Yutong Bai, Changhu Wang, Alan Yuille).

Table Extraction Tool

Semantic-based Patch Detection for Binary Programs

Distort a video using Seam Carving (video) and Vibrato effect (sound)

This is a c++ project deploying a deep scene text reading pipeline with tensorflow. It reads text from natural scene images. It uses frozen tensorflow graphs. The detector detect scene text locations. The recognizer reads word from each detected bounding box.

CNN+LSTM+CTC based OCR implemented using tensorflow.

Code for the head detector (HeadHunter) proposed in our CVPR 2021 paper Tracking Pedestrian Heads in Dense Crowd.

This is a real life mario project using python and mediapipe

Code for CVPR 2022 paper "SoftGroup for Instance Segmentation on 3D Point Clouds"

7th place solution

Morphological edge detection or object's boundary detection using erosion and dialation in OpenCV python

Textboxes implementation with Tensorflow (python)

A novel region proposal network for more general object detection ( including scene text detection ).

Introduction to Augmented Reality (AR) with Python 3 and OpenCV 4.2.

code for our ICCV 2021 paper "DeepCAD: A Deep Generative Network for Computer-Aided Design Models"

Generates a message from the infamous Jerma Impostor image

The first open-source library that detects the font of a text in a image.