Cross-view Transformers for real-time Map-view Semantic Segmentation (CVPR 2022 Oral)

Last update: Dec 25, 2022

Overview

Cross View Transformers

This repository contains the source code and data for our paper:

Cross-view Transformers for real-time Map-view Semantic Segmentation
Brady Zhou, Philipp Krähenbühl
CVPR 2022

Demos

Map-view Segmentation: The model uses multi-view images to produce a map-view segmentation at 45 FPS

Map Making: With vehicle pose, we can construct a map by fusing model predictions over time

Cross-view Attention: For a given map-view location, we show which image patches are being attended to

Installation

# Clone repo
git clone https://github.com/bradyz/cross_view_transformers.git

cd cross_view_transformers

# Setup conda environment
conda create -y --name cvt python=3.8

conda activate cvt
conda install -y pytorch torchvision cudatoolkit=11.3 -c pytorch

# Install dependencies
pip install -r requirements.txt
pip install -e .

Data

Documentation:

Dataset setup
Label generation (optional)

Download the original datasets and our generated map-view labels

	Dataset	Labels
nuScenes	keyframes + map expansion (60 GB)	cvt_labels_nuscenes.tar.gz (361 MB)
Argoverse 1.1	3D tracking	coming soon™

The structure of the extracted data should look like the following

/datasets/
├─ nuscenes/
│  ├─ v1.0-trainval/
│  ├─ v1.0-mini/
│  ├─ samples/
│  ├─ sweeps/
│  └─ maps/
│     ├─ basemap/
│     └─ expansion/
└─ cvt_labels_nuscenes/
   ├─ scene-0001/
   ├─ scene-0001.json
   ├─ ...
   ├─ scene-1000/
   └─ scene-1000.json

When everything is setup correctly, check out the dataset with

python3 scripts/view_data.py \
  data=nuscenes \
  data.dataset_dir=/media/datasets/nuscenes \
  data.labels_dir=/media/datasets/cvt_labels_nuscenes \
  data.version=v1.0-mini \
  visualization=nuscenes_viz \
  +split=val

Training

An average job of 50k training iterations takes ~8 hours.
Our models were trained using 4 GPU jobs, but also can be trained on single GPU.

To train a model,

python3 scripts/train.py \
  +experiment=cvt_nuscenes_vehicle
  data.dataset_dir=/media/datasets/nuscenes \
  data.labels_dir=/media/datasets/cvt_labels_nuscenes

For more information, see

config/config.yaml - base config
config/model/cvt.yaml - model architecture
config/experiment/cvt_nuscenes_vehicle.yaml - additional overrides

Additional Information

Awesome Related Repos

License

This project is released under the MIT license

Citation

If you find this project useful for your research, please use the following BibTeX entry.

@inproceedings{zhou2022cross,
    title={Cross-view Transformers for real-time Map-view Semantic Segmentation},
    author={Zhou, Brady and Kr{\"a}henb{\"u}hl, Philipp},
    booktitle={CVPR},
    year={2022}
}

Cross-view Transformers for real-time Map-view Semantic Segmentation (CVPR 2022 Oral)

Related tags

Overview

Cross View Transformers

Demos

Installation

Data

Training

Additional Information

Awesome Related Repos

License

Citation

Owner

Brady Zhou

Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis Implementation

Revisiting Discriminator in GAN Compression: A Generator-discriminator Cooperative Compression Scheme (NeurIPS2021)

aka "Bayesian Methods for Hackers": An introduction to Bayesian methods + probabilistic programming with a computation/understanding-first, mathematics-second point of view. All in pure Python ;)

NOMAD - A blackbox optimization software

Implementation of Convolutional LSTM in PyTorch.

Official PyTorch implementation of BlobGAN: Spatially Disentangled Scene Representations

Neural Network to colorize grayscale images

LaneDet is an open source lane detection toolbox based on PyTorch that aims to pull together a wide variety of state-of-the-art lane detection models

The implementation of PEMP in paper "Prior-Enhanced Few-Shot Segmentation with Meta-Prototypes"

The official implementation of Autoregressive Image Generation using Residual Quantization (CVPR '22)

Official repository with code and data accompanying the NAACL 2021 paper "Hurdles to Progress in Long-form Question Answering" (https://arxiv.org/abs/2103.06332).

Repository to run object detection on a model trained on an autonomous driving dataset.

Finetuning Pipeline

Implementation of [Time in a Box: Advancing Knowledge Graph Completion with Temporal Scopes].

Lightwood is Legos for Machine Learning.

Faune proche - Retrieval of Faune-France data near a google maps location

An AI made using artificial intelligence (AI) and machine learning algorithms (ML) .

Easy-to-use,Modular and Extendible package of deep-learning based CTR models .

Shape Matching of Real 3D Object Data to Synthetic 3D CADs (3DV project @ ETHZ)

Spectral Tensor Train Parameterization of Deep Learning Layers