Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

Last update: Nov 07, 2022

Related tags

Overview

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

We implement the classification, object detection and instance segmentation tasks based on our cvpods. The users should install cvpods first and run the experiments in this repo.

Changelog

4.23.2021 Update the DisAlign on LVIS v0.5(Mask R-CNN + Res50)
4.12.2021 Update the README

0. How to Use

Step-1: Install the latest cvpods.
Step-2: cd cvpods
Step-3: Prepare dataset for different tasks.
Step-4: git clone https://github.com/Megvii-BaseDetection/DisAlign playground_disalign
Step-5: Enter one folder and run pods_train --num-gpus 8
Step-6: Use pods_test --num-gpus 8 to evaluate the last the checkpoint

1. Image Classification

We support the the following three datasets:

ImageNet-LT Dataset
iNaturalist-2018 Dataset
Place-LT Dataset

We refer the user to CLS_README for more details.

2. Object Detection/Instance Segmentation

We support the two versions of the LVIS dataset:

LVIS v0.5
LVIS v1.0

Highlight

To speedup the evaluation on LVIS dataset, we provide the C++ optimized evaluation api by modifying the coco_eval(C++) in cvpods.

The C++ version lvis_eval API will save ~30% time when calculating the mAP.

We provide support for the metric of AP_fixed and AP_pool proposed in large-vocab-devil
We will support more recent works on long-tail detection in this project(e.g. EQLv2, CenterNet2, etc.) in the future.

We refer the user to DET_README for more details.

3. Semantic Segmentation

We adopt the mmsegmentation as the codebase for runing all experiments of DisAlign. Currently, the user should use DisAlign_Seg for the semantic segmentation experiments. We will add the support for these experiments in cvpods in the future.

Acknowledgement

Thanks for the following projects:

Citing DisAlign

If you are using the DisAlign in your research or with to refer to the baseline results publised in this repo, please use the following BibTex entry.

@inproceedings{zhang2021disalign,
  title={Distribution Alignment: A Unified Framework for Long-tail Visual Recognition.},
  author={Zhang, Songyang and Li, Zeming and Yan, Shipeng and He, Xuming and Sun, Jian},
  booktitle={CVPR},
  year={2021}
}

License

This repo is released under the Apache 2.0 license. Please see the LICENSE file for more information.

Comments

scale in cosine classifier

Hi, thanks for your great work! I notice you use the cosine classifier in many experiments and it can get a better baseline. The formula is as follows

I am wondering the value of s?

opened by L1aoXingyu 5
Is it correct to freeze the weight and bias of the DisAlign Linear Layer as well?

Hello. Thank you for your project! I'm testing your code on my custom dataset. My task is classification. I have a question about your code implementation.

https://github.com/Megvii-BaseDetection/DisAlign/blob/a2fc3500a108cb83e3942293a5675c97ab3a2c6e/classification/imagenetlt/resnext50/resx50.scratch.imagenet_lt.224size.90e.disalign.10e/net.py#L56-L62

From my understanding, in stage 2, remove the linear layer used in stage 1 and add DisAlign Linear Layer. And freeze all parts except for logit_scale, logit_bias, and confidence_layer. At this time, the weight and bias of DisAlignLinear are also frozen. (self.weight, self.bias) Is my understanding correct?

If so, are the weight and bias of DisAlignLinearLayer fixed after the initialization? (The weight and bias of the linear layer in stage 1 are not copied either)

If my understanding is correct, why is the weight of DisAlignLinear also frozen?

I will wait for your reply. thanks!

opened by jeongHwarr 4
Where is the DisAlignLinear module?

Hello. Thank you for your impressive project!

I want to apply DisAlign to classification. However, an error occurs in the import part. https://github.com/Megvii-BaseDetection/DisAlign/blob/a2fc3500a108cb83e3942293a5675c97ab3a2c6e/classification/imagenetlt/resnext50/resx50.scratch.imagenet_lt.224size.90e.disalign.10e/net.py#L7 I coudn't find the DisAlignLinear in cvpods.layers. and there also isn't exist at https://github.com/Megvii-BaseDetection/cvpods/tree/master/cvpods/layers How can I solve this problem?

Thank you!

opened by jeongHwarr 4
Can someone kindly share their codes of Classification task on ImageNet_LT?

I tried to train the proposed method on ImageNet_LT, but I can only get an average testing rate about 49%, which is far from the rate described in the paper (52.9). Some of the details regarding my implementations are given as follows: (1) The feature extractor is ResNexT-50 and the head classifier is a linear classifier. The testing accuracy in Stage-One is 43.9%, which is OK.

(2) The testing accuracy of adopting cRT method in Stage-Two is 49.6%, which is identical to one reported in other papers. (3) When fine-tuning the model in Stage-2, both the feature-extractor and head classifier are frozen, and a DisAliLinear model (which is implemented in CVPODs) is retrained. The testing accuracy can only reach 48.8%, which is far away from the one reported in your paper.

opened by smallcube 4
The code for semantic segmentation is missing

Hi, thank you for the nice work, but the code for semantic segmentation is missing and the URL for it in the README could not be opened. Could you please fix this issue?

opened by curiosity654 3
About the reference Distribution p_r in Eq. (10)

Hi, Thank you for providing your code. Here I was wondering the Equation (10) in your paper (The definition of p_r), which seems not to be a distribution. Since every x_i can only have one label, the reference distribution p_r(y| x_i) will be the distribution like (0, 0, 0,...,w_c, 0, 0,...,0). And the sum of this distribution is w_c, but not 1.

Could you help me understand this equation? Thanks in advance.

opened by Kevinz-code 3
import error

Hi, thanks for the great work. Maybe I missed it, but it seems that the code for this project has been incorporated into cvpods. I couldn't launch any experiments due to ImportErrors like: from cvpods.layers import DisAlignLinear ImportError: cannot import name 'DisAlignLinear' from 'cvpods.layers' Also, I didn't find the corresponding functions in cvpods.

Any help will be appreciated. Thanks.

opened by YUE-FAN 2
about the confidence score σ(x)

In the paper, the σ(x) is implemented as a linear layer followed by a non-linear activation function (e.g., sigmoid function) for all input x. How to understand the input x？the matrix of raw iamge, or the extracted features, even or cls_score? Thank you!

opened by lzed2399 2
exp_reweight = exp_reweight / np.sum(exp_reweight) * num_foreground
Dear author, I have some questions about the code and paper:

exp_reweight = exp_reweight / np.sum(exp_reweight) * num_foreground Why "exp_reweight" is multiplied by the coefficient "num_foreground"? It is not mentioned in the paper.

Is "K" in the empirical class frequencies r = [r1, · · · , rK] on the training set in the paper the same as the class number C of the training set?
opened by Liu-wanbing 2
The DisAlign_Seg page can't open

Hi, thank you for your amazing work. I want to have a try on semantic segmentation problem. But the link https://github.com/Megvii-BaseDetection/DisAlign/blob/main/TODO of DisAlign_Seg cannot open. Could you pls have a look? Thank you.

opened by Kittywyk 1
Do you use validation dataset?

https://github.com/Megvii-BaseDetection/DisAlign/blob/main/classification/imagenetlt/resnext50/resx50.scratch.imagenet_lt.224size.90e.disalign.10e/config.py#L31

It seems that you only use test dataset? What is the reason for that?

opened by qianlanwyd 1
How can I test and augtest the trained semseg DisAlign model?

Thank you for your great work. I met the question of 'AttributeError: 'ConfigDict' object has no attribute 'model' in testing the trained semseg DisAlign model when running the file https://github.com/Megvii-BaseDetection/DisAlign/blob/main/semantic_seg/exps/ade20k_fcn_disalign/disalign_fcn_r50-d8_512x512_160k_ade20k.sh. Is this a version problem of mmcv-full or something else? And can you perfect the explanation of the disalign part of Readme.md in semantic_seg?

opened by jh151170 0
the code question in semantic_seg

Hi, I have a questation about the logit_scale and logit_bias in semantic_seg. The shape of the above parameter is (1, num_classes, 1, 1), why not is (1, num_classes, 512, 512) which is matched the input image size for semantic segmenation.

opened by Ianresearch 8
Value of the learned scale and bias vector?

Hi, did you check the value change of the learned scale and bias vector throughout the training process? I find the value of them change in the first few iterations and remain stable in the rest time on my own classification dataset. I wonder how the learned vectors look like in your paper? Thanks!

opened by Jacobew 1

Releases(LVIS)

LVIS(May 7, 2021)

Source code(tar.gz)
Source code(zip)
imagenet_lt_category_frequency.json(9.20 KB)
num_shots_v0.5.npy(9.73 KB)
num_shots_v1.0.json(19.85 KB)
resx50.scratch.imagenet_lt.224size.90e.disalign.10e.model_final_plain.pth(95.86 MB)
resx50.scratch.imagenet_lt.224size.90e.model_final_plain.pth(95.84 MB)

Owner

BaseDetection Team of Megvii

GitHub Repository

ONNX Runtime Web demo is an interactive demo portal showing real use cases running ONNX Runtime Web in VueJS.

ONNX Runtime Web demo is an interactive demo portal showing real use cases running ONNX Runtime Web in VueJS. It currently supports four examples for you to quickly experience the power of ONNX Runti

58 Dec 18, 2022

Deep Ensemble Learning with Jet-Like architecture

Ransomware analysis using DEL with jet-like architecture comprising two CNN wings, a sparse AE tail, a non-linear PCA to produce a diverse feature space, and an MLP nose

2 Feb 06, 2022

StrongSORT: Make DeepSORT Great Again

StrongSORT StrongSORT: Make DeepSORT Great Again StrongSORT: Make DeepSORT Great Again Yunhao Du, Yang Song, Bo Yang, Yanyun Zhao arxiv 2202.13514 Abs

369 Jan 04, 2023

Pytorch GUI(demo) for iVOS(interactive VOS) and GIS (Guided iVOS)

GUI for iVOS(interactive VOS) and GIS (Guided iVOS) GUI Implementation of CVPR2021 paper "Guided Interactive Video Object Segmentation Using Reliabili

13 Dec 09, 2022

Directed Greybox Fuzzing with AFL

AFLGo: Directed Greybox Fuzzing AFLGo is an extension of American Fuzzy Lop (AFL). Given a set of target locations (e.g., folder/file.c:582), AFLGo ge

380 Nov 24, 2022

A code implementation of AC-GC: Activation Compression with Guaranteed Convergence, in NeurIPS 2021.

Code For AC-GC: Lossy Activation Compression with Guaranteed Convergence This code is intended to be used as a supplemental material for submission to

2 Nov 01, 2022

Human4D Dataset tools for processing and visualization

HUMAN4D: A Human-Centric Multimodal Dataset for Motions & Immersive Media HUMAN4D constitutes a large and multimodal 4D dataset that contains a variet

15 Nov 09, 2022

[Link]mareteutral - pars tradg wth M []

pairs-trading-with-ML Jonathan Larkin, August 2017 One popular strategy classification is Pairs Trading. Though this category of strategies can exhibi

134 Jan 06, 2023

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL)

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL) This repository contains all source code used to generate the results in the article "

3 Jul 23, 2022

“英特尔创新大师杯”深度学习挑战赛赛道3：CCKS2021中文NLP地址相关性任务

ccks2021-track3 CCKS2021中文NLP地址相关性任务-赛道三-冠军方案团队：我的加菲鱼- wodejiafeiyu 初赛第二/复赛第一/决赛第一前言 19年开始，陆陆续续参加了一些比赛，拿到过一些top，比较懒一直都没分享过，这次比较幸运又拿了top1，打算分享下分类的任务

131 Dec 31, 2022

pix2pix in tensorflow.js

pix2pix in tensorflow.js This repo is moved to https://github.com/yining1023/pix2pix_tensorflowjs_lite See a live demo here: https://yining1023.github

47 Oct 04, 2022

Improving Query Representations for DenseRetrieval with Pseudo Relevance Feedback:A Reproducibility Study.

APR The repo for the paper Improving Query Representations for DenseRetrieval with Pseudo Relevance Feedback:A Reproducibility Study. Environment setu

8 Nov 26, 2022

这个开源项目主要是对经典的时间序列预测算法论文进行复现，模型主要参考自GluonTS，框架主要参考自Informer

Time Series Research with Torch 这个开源项目主要是对经典的时间序列预测算法论文进行复现，模型主要参考自GluonTS，框架主要参考自Informer。建立原因相较于mxnet和TF，Torch框架中的神经网络层需要提前指定输入维度： # 建立线性层 TensorF

85 Dec 29, 2022

Very large and sparse networks appear often in the wild and present unique algorithmic opportunities and challenges for the practitioner

Sparse network learning with snlpy Very large and sparse networks appear often in the wild and present unique algorithmic opportunities and challenges

1 Apr 30, 2021

Reinforcement learning framework and algorithms implemented in PyTorch.

2.1k Jan 04, 2023

Hunt down social media accounts by username across social networks

Hunt down social media accounts by username across social networks Installation | Usage | Docker Notes | Contributing Installation # clone the repo $

1 Dec 14, 2021

Implemented fully documented Particle Swarm Optimization algorithm (basic model with few advanced features) using Python programming language

Implemented fully documented Particle Swarm Optimization (PSO) algorithm in Python which includes a basic model along with few advanced features such as updating inertia weight, cognitive, social lea

9 Nov 29, 2022

Implementation of "Distribution Alignment: A Unified Framework for Long-tail Visual Recognition"(CVPR 2021)

Related tags

Overview

Changelog

0. How to Use

1. Image Classification

2. Object Detection/Instance Segmentation

3. Semantic Segmentation

Acknowledgement

Citing DisAlign

License

Comments

Releases(LVIS)

LVIS(May 7, 2021)

Owner

ONNX Runtime Web demo is an interactive demo portal showing real use cases running ONNX Runtime Web in VueJS.

Deep Ensemble Learning with Jet-Like architecture

StrongSORT: Make DeepSORT Great Again

Pytorch GUI(demo) for iVOS(interactive VOS) and GIS (Guided iVOS)

Directed Greybox Fuzzing with AFL

A code implementation of AC-GC: Activation Compression with Guaranteed Convergence, in NeurIPS 2021.

Human4D Dataset tools for processing and visualization

[Link]mareteutral - pars tradg wth M []

Surrogate- and Invariance-Boosted Contrastive Learning (SIB-CL)

“英特尔创新大师杯”深度学习挑战赛 赛道3：CCKS2021中文NLP地址相关性任务

pix2pix in tensorflow.js

Improving Query Representations for DenseRetrieval with Pseudo Relevance Feedback:A Reproducibility Study.

这个开源项目主要是对经典的时间序列预测算法论文进行复现，模型主要参考自GluonTS，框架主要参考自Informer

Very large and sparse networks appear often in the wild and present unique algorithmic opportunities and challenges for the practitioner

Reinforcement learning framework and algorithms implemented in PyTorch.

Hunt down social media accounts by username across social networks

Implemented fully documented Particle Swarm Optimization algorithm (basic model with few advanced features) using Python programming language

Unicorn can be used for performance analyses of highly configurable systems with causal reasoning

The code of NeurIPS 2021 paper "Scalable Rule-Based Representation Learning for Interpretable Classification".

Personalized Transfer of User Preferences for Cross-domain Recommendation (PTUPCDR)

“英特尔创新大师杯”深度学习挑战赛赛道3：CCKS2021中文NLP地址相关性任务