NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework

Last update: Dec 08, 2022

Related tags

Overview

NLP From Scratch Without Large-Scale Pretraining

This repository contains the code, pre-trained model checkpoints and curated datasets for our paper: NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework.

In our proposed framework, named TLM (task-driven language modeling), instead of training a language model over the entire general corpus and then finetuning it on task data, we first usetask data as queries to retrieve a tiny subset of the general corpus, and then perform joint learning on both the task objective and self-supervised language modeling objective.

Requirements

We implement our models and training loops based on the opensource products from HuggingFace. The core denpencies of this repository are listed in requirements.txt, which can be installed through:

pip install -r requirements.txt

All our experiments are conducted on a node with 8 A100 40GB SXM gpus. Different computational devices may result slightly different results from the reported ones.

Models and Datasets

We release the trained models on 8 tasks with 3 different scales, together with the task datasets and selected external data. Our released model checkpoints, datasets and the performance of each model for each task are listed in the following table.

	AGNews	Hyp.	Help.	IMDB	ACL.	SciERC	Chem.	RCT
Small	93.74	93.53	70.54	93.08	69.84	80.51	81.99	86.99
Medium	93.96	94.05	70.90	93.97	72.37	81.88	83.24	87.28
Large	94.36	95.16	72.49	95.77	72.19	83.29	85.12	87.50

The released models and datasets are compatible with HuggingFace's Transformers and Datasets. We provide an example script to evaluate a model checkpoints on a certain task, run

bash example_scripts/evaluate.sh

To get the evaluation results for SciERC with a small-scale model.

Training

We provide two example scripts to train a model from scratch, run

bash example_scripts/train.sh && bash example_scripts/finetune.sh

To train a small-scale model for SciERC. Here example_scripts/train.sh corresponds to the first stage training where the external data ratio and MLM weight are non-zero, and example_scripts/finetune.sh corresponds to the second training stage where no external data or self-supervised loss can be perceived by the model.

Citation

Please cite our paper if you use TLM in your work:

@misc{yao2021tlm,
title={NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework},
author={Yao, Xingcheng and Zheng, Yanan and Yang, Xiaocong and Yang, Zhilin},
year={2021}
}

NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework

Related tags

Overview

NLP From Scratch Without Large-Scale Pretraining

Requirements

Models and Datasets

Training

Citation

Owner

Xingcheng Yao

A lightweight face-recognition toolbox and pipeline based on tensorflow-lite

A Pytorch Implementation of Domain adaptation of object detector using scissor-like networks

Voxel Transformer for 3D object detection

[CVPRW 21] "BNN - BN = ? Training Binary Neural Networks without Batch Normalization", Tianlong Chen, Zhenyu Zhang, Xu Ouyang, Zechun Liu, Zhiqiang Shen, Zhangyang Wang

An official implementation of the Anchor DETR.

Generative Modelling of BRDF Textures from Flash Images [SIGGRAPH Asia, 2021]

Official code of the paper "ReDet: A Rotation-equivariant Detector for Aerial Object Detection" (CVPR 2021)

WRENCH: Weak supeRvision bENCHmark

Raster Vision is an open source Python framework for building computer vision models on satellite, aerial, and other large imagery sets

Manipulation OpenAI Gym environments to simulate robots at the STARS lab

Accommodating supervised learning algorithms for the historical prices of the world's favorite cryptocurrency and boosting it through LightGBM.

OpenMMLab Video Perception Toolbox. It supports Video Object Detection (VID), Multiple Object Tracking (MOT), Single Object Tracking (SOT), Video Instance Segmentation (VIS) with a unified framework.

ECLARE: Extreme Classification with Label Graph Correlations

PyTorch implementation of the paper Dynamic Token Normalization Improves Vision Transfromers.

SPT_LSA_ViT - Implementation for Visual Transformer for Small-size Datasets

InsightFace: 2D and 3D Face Analysis Project on MXNet and PyTorch

MultiLexNorm 2021 competition system from ÚFAL

PyTorch reimplementation of the Smooth ReLU activation function proposed in the paper "Real World Large Scale Recommendation Systems Reproducibility and Smooth Activations" [arXiv 2022].

[ICLR 2021 Spotlight Oral] "Undistillable: Making A Nasty Teacher That CANNOT teach students", Haoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You, Xiaohui Xie, Zhangyang Wang

Sparse-dense operators implementation for Paddle