Pytorch implemenation of Stochastic Multi-Label Image-to-image Translation (SMIT)

Last update: Mar 01, 2022

Overview

SMIT: Stochastic Multi-Label Image-to-image Translation

This repository provides a PyTorch implementation of SMIT. SMIT can stochastically translate an input image to multiple domains using only a single generator and a discriminator. It only needs a target domain (binary vector e.g., [0,1,0,1,1] for 5 different domains) and a random gaussian noise.

Paper

SMIT: Stochastic Multi-Label Image-to-image Translation
Andrés Romero¹, Pablo Arbelaez¹, Luc Van Gool², Radu Timofte²
¹Biomedical Computer Vision (BCV) Lab, Universidad de Los Andes.
²Computer Vision Lab (CVL), ETH Zürich.

Citation

@article{romero2019smit,
  title={SMIT: Stochastic Multi-Label Image-to-Image Translation},
  author={Romero, Andr{\'e}s and Arbel{\'a}ez, Pablo and Van Gool, Luc and Timofte, Radu},
  journal={ICCV Workshops},
  year={2019}
}

Dependencies

Python (2.7, 3.5+)
PyTorch (0.3, 0.4, 1.0)

Usage

Cloning the repository

$ git clone https://github.com/BCV-Uniandes/SMIT.git
$ cd SMIT

Downloading the dataset

To download the CelebA dataset:

$ bash generate_data/download.sh

Train command:

./main.py --GPU=$gpu_id --dataset_fake=CelebA

Each dataset must has datasets/ .py and datasets/ .yaml files. All models and figures will be stored at snapshot/models/$dataset_fake/ _ .pth and snapshot/samples/$dataset_fake/ _ .jpg, respectivelly.

Test command:

./main.py --GPU=$gpu_id --dataset_fake=CelebA --mode=test

SMIT will expect the .pth weights are stored at snapshot/models/$dataset_fake/ (or --pretrained_model=location/model.pth should be provided). If there are several models, it will take the last alphabetical one.

Demo:

./main.py --GPU=$gpu_id --dataset_fake=CelebA --mode=test --DEMO_PATH=location/image_jpg/or/location/dir

DEMO performs transformation per attribute, that is swapping attributes with respect to the original input as in the images below. Therefore, --DEMO_LABEL is provided for the real attribute if DEMO_PATH is an image (If it is not provided, the discriminator acts as classifier for the real attributes).

Pretrained models

Models trained using Pytorch 1.0.

Multi-GPU

For multiple GPUs we use Horovod. Example for training with 4 GPUs:

mpirun -n 4 ./main.py --dataset_fake=CelebA

Qualitative Results. Multi-Domain Continuous Interpolation.

First column (original input) -> Last column (Opposite attributes: smile, age, genre, sunglasses, bangs, color hair). Up: Continuous interpolation for the fake image. Down: Continuous interpolation for the attention mechanism.

Pytorch implemenation of Stochastic Multi-Label Image-to-image Translation (SMIT)

Related tags

Overview

SMIT: Stochastic Multi-Label Image-to-image Translation

Paper

Citation

Dependencies

Usage

Cloning the repository

Downloading the dataset

Train command:

Test command:

Demo:

Pretrained models

Multi-GPU

Qualitative Results. Multi-Domain Continuous Interpolation.

Qualitative Results. Random sampling.

CelebA

EmotionNet

RafD

Edges2Shoes

Edges2Handbags

Yosemite

Painters

Qualitative Results. Style Interpolation between first and last row.

CelebA

EmotionNet

RafD

Edges2Shoes

Edges2Handbags

Yosemite

Painters

Qualitative Results. Label continuous inference between first and last row.

CelebA

EmotionNet

Owner

Biomedical Computer Vision Group @ Uniandes

Tools for manipulating UVs in the Blender viewport.

Running AlphaFold2 (from ColabFold) in Azure Machine Learning

git《Self-Attention Attribution: Interpreting Information Interactions Inside Transformer》(AAAI 2021) GitHub:

This is the official repository of Music Playlist Title Generation: A Machine-Translation Approach.

RoMA: Robust Model Adaptation for Offline Model-based Optimization

Uses Open AI Gym environment to create autonomous cryptocurrency bot to trade cryptocurrencies.

Contrastive Feature Loss for Image Prediction

[v1 (ISBI'21) + v2] MedMNIST: A Large-Scale Lightweight Benchmark for 2D and 3D Biomedical Image Classification

Contains code for Deep Kernelized Dense Geometric Matching

"Neural Turing Machine" in Tensorflow

Improving the robustness and performance of biomedical NLP models through adversarial training

ImVoxelNet: Image to Voxels Projection for Monocular and Multi-View General-Purpose 3D Object Detection

A python library for face detection and features extraction based on mediapipe library

PyTorch Code of "Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics"

DirectVoxGO reconstructs a scene representation from a set of calibrated images capturing the scene.

Official Pytorch Implementation of: "Semantic Diversity Learning for Zero-Shot Multi-label Classification"(2021) paper

Multi-Horizon-Forecasting-for-Limit-Order-Books

A Kaggle competition: discriminate gender based on handwriting

[NeurIPS 2020] This project provides a strong single-stage baseline for Long-Tailed Classification, Detection, and Instance Segmentation (LVIS).

Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution Networks (MAPDN)