Codebase for ECCV18 "The Sound of Pixels"

Last update: Dec 20, 2022

Overview

Sound-of-Pixels

Codebase for ECCV18 "The Sound of Pixels".

*This repository is under construction, but the core parts are already there.

Environment

The code is developed under the following configurations.

Hardware: 1-4 GPUs (change [--num_gpus NUM_GPUS] accordingly)
Software: Ubuntu 16.04.3 LTS, CUDA>=8.0, Python>=3.5, PyTorch>=0.4.0

Training

Prepare video dataset.

a. Download MUSIC dataset from: https://github.com/roudimit/MUSIC_dataset

b. Download videos.

Preprocess videos. You can do it in your own way as long as the index files are similar.

a. Extract frames at 8fps and waveforms at 11025Hz from videos. We have following directory structure:

data
├── audio
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp3
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp3
│   |   ├── ...
│   ├── ...
|
└── frames
|   ├── acoustic_guitar
│   |   ├── M3dekVSwNjY.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── trumpet
│   |   ├── STKXyBGSGyE.mp4
│   |   |   ├── 000001.jpg
│   |   |   ├── ...
│   |   ├── ...
│   ├── ...

b. Make training/validation index files by running:

python scripts/create_index_files.py

It will create index files train.csv/val.csv with the following format:

./data/audio/acoustic_guitar/M3dekVSwNjY.mp3,./data/frames/acoustic_guitar/M3dekVSwNjY.mp4,1580
./data/audio/trumpet/STKXyBGSGyE.mp3,./data/frames/trumpet/STKXyBGSGyE.mp4,493

For each row, it stores the information: AUDIO_PATH,FRAMES_PATH,NUMBER_FRAMES

Train the default model.

./scripts/train_MUSIC.sh

During training, visualizations are saved in HTML format under ckpt/MODEL_ID/visualization/.

Evaluation

(Optional) Download our trained model weights for evaluation.

./scripts/download_trained_model.sh

Evaluate the trained model performance.

./scripts/eval_MUSIC.sh

Reference

If you use the code or dataset from the project, please cite:

    @InProceedings{Zhao_2018_ECCV,
        author = {Zhao, Hang and Gan, Chuang and Rouditchenko, Andrew and Vondrick, Carl and McDermott, Josh and Torralba, Antonio},
        title = {The Sound of Pixels},
        booktitle = {The European Conference on Computer Vision (ECCV)},
        month = {September},
        year = {2018}
    }

Codebase for ECCV18 "The Sound of Pixels"

Related tags

Overview

Sound-of-Pixels

Environment

Training

Evaluation

Reference

Owner

Hang Zhao

Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network

MemStream: Memory-Based Anomaly Detection in Multi-Aspect Streams with Concept Drift

Official implementation of the MM'21 paper Constrained Graphic Layout Generation via Latent Optimization

Deep Learning as a Cloud API Service.

The official repository for BaMBNet

Cartoon-StyleGan2 🙃 : Fine-tuning StyleGAN2 for Cartoon Face Generation

An unopinionated replacement for PyTorch's Dataset and ImageFolder, that handles Tar archives

The InterScript dataset contains interactive user feedback on scripts generated by a T5-XXL model.

Code for ICE-BeeM paper - NeurIPS 2020

This repository provides code for "On Interaction Between Augmentations and Corruptions in Natural Corruption Robustness".

A toolkit for document-level event extraction, containing some SOTA model implementations

Python/Rust implementations and notes from Proofs Arguments and Zero Knowledge

This repo is a C++ version of yolov5_deepsort_tensorrt. Packing all C++ programs into .so files, using Python script to call C++ programs further.

A short code in python, Enchpyter, is able to encrypt and decrypt words as you determine, of course

Final report with code for KAIST Course KSE 801.

A forwarding MPI implementation that can use any other MPI implementation via an MPI ABI

nn_builder lets you build neural networks with less boilerplate code

HyperaPy: An automatic hyperparameter optimization framework ⚡🚀

Code release for Universal Domain Adaptation(CVPR 2019)

"Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices", official implementation