Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Last update: Nov 25, 2022

Related tags

Deep Learning SGN

Overview

Semantic Grouping Network for Video Captioning

Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. AAAI 2021. [arxiv]

Environment

Ubuntu 16.04
CUDA 9.2
cuDNN 7.4.2
Java 8
Python 2.7.12
- PyTorch 1.1.0
- Other python packages specified in requirements.txt

Usage

1. Setup

$ pip install -r requirements.txt

2. Prepare Data

Download the GloVe Embedding from here and locate it at data/Embeddings/GloVe/GloVe_300.json.
Extract features from datasets and locate them at data/ /features/ .hdf5.

e.g. ResNet101 features of the MSVD dataset will be located at data/MSVD/features/ResNet101.hdf5.

I refer to this repo for extracting the ResNet101 features, and this repo for extracting the 3D-ResNext101 features.
Split the features into train, val, and test sets by running following commands.
```
$ python -m split.MSVD
$ python -m split.MSR-VTT
```

You can skip step 2-3 and download below files

MSVD
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]
MSR-VTT
- ResNet-101 [train] [val] [test]
- 3D-ResNext-101 [train] [val] [test]

3. Prepare The Code for Evaluation

Clone the evaluation code from the official coco-evaluation repo.

$ git clone https://github.com/tylin/coco-caption.git
$ mv coco-caption/pycocoevalcap .
$ rm -rf coco-caption

4. Extract Negative Videos

$ python extract_negative_videos.py

or you can skip this step as the output files are already uploaded at data/ /metadata/neg_vids_ .json

5. Train

$ python train.py

You can change some hyperparameters by modifying config.py.

Pretrained Models - SGN(R101+RN)

*Disclaimer: The models above do not have the same weight as the models used in the paper (I trained them again because I lost).

6. Evaluate

$ python evaluate.py --ckpt_fpath

License

The source-code in this repository is released under MIT License.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Get started with Machine Learning with Python - An introduction with Python programming examples

WaveFake: A Data Set to Facilitate Audio DeepFake Detection

3D cascade RCNN for object detection on point cloud

Global Pooling, More than Meets the Eye: Position Information is Encoded Channel-Wise in CNNs, ICCV 2021

Generate Contextual Directory Wordlist For Target Org

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Neural Scene Flow Prior (NeurIPS 2021 spotlight)

DLFlow is a deep learning framework.

Implementation for Stankevičiūtė et al. "Conformal time-series forecasting", NeurIPS 2021.

🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022

A PyTorch implementation of Implicit Q-Learning

A program that uses computer vision to detect hand gestures, used for controlling movie players.

Code and datasets for the paper "KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction"

Official PyTorch implementation of the paper "Recycling Discriminator: Towards Opinion-Unaware Image Quality Assessment Using Wasserstein GAN", accepted to ACM MM 2021 BNI Track.

InferPy: Deep Probabilistic Modeling with Tensorflow Made Easy

This is the Pytorch implementation of Progressive Attentional Manifold Alignment.

A Deep Learning Based Knowledge Extraction Toolkit for Knowledge Base Population

Implementation of paper "Self-supervised Learning on Graphs:Deep Insights and New Directions"

a minimal terminal with python 😎😉

CS550 Machine Learning course project on CNN Detection.

Official pytorch implementation of the AAAI 2021 paper Semantic Grouping Network for Video Captioning

Related tags

Overview

Semantic Grouping Network for Video Captioning

Environment

Usage

1. Setup

2. Prepare Data

3. Prepare The Code for Evaluation

4. Extract Negative Videos

5. Train

6. Evaluate

License

Owner

Hobin Ryu

Get started with Machine Learning with Python - An introduction with Python programming examples

WaveFake: A Data Set to Facilitate Audio DeepFake Detection

3D cascade RCNN for object detection on point cloud

Global Pooling, More than Meets the Eye: Position Information is Encoded Channel-Wise in CNNs, ICCV 2021

Generate Contextual Directory Wordlist For Target Org

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Neural Scene Flow Prior (NeurIPS 2021 spotlight)

DLFlow is a deep learning framework.

Implementation for Stankevičiūtė et al. "Conformal time-series forecasting", NeurIPS 2021.

🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022

A PyTorch implementation of Implicit Q-Learning

A program that uses computer vision to detect hand gestures, used for controlling movie players.

Code and datasets for the paper "KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction"

Official PyTorch implementation of the paper "Recycling Discriminator: Towards Opinion-Unaware Image Quality Assessment Using Wasserstein GAN", accepted to ACM MM 2021 BNI Track.

InferPy: Deep Probabilistic Modeling with Tensorflow Made Easy

​ This is the Pytorch implementation of Progressive Attentional Manifold Alignment.

A Deep Learning Based Knowledge Extraction Toolkit for Knowledge Base Population

Implementation of paper "Self-supervised Learning on Graphs:Deep Insights and New Directions"

a minimal terminal with python 😎😉

CS550 Machine Learning course project on CNN Detection.

This is the Pytorch implementation of Progressive Attentional Manifold Alignment.