Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences

Last update: Dec 20, 2022

Overview

Overview

Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences. Copula and functional Principle Component Analysis (fPCA) are statistical models that allow these properties to be simulated (Joe 2014). As such, copula generated data have shown potential to improve the generalization of machine learning (ML) emulators (Meyer et al. 2021) or anonymize real-data datasets (Patki et al. 2016).

Synthia is an open source Python package to model univariate and multivariate data, parameterize data using empirical and parametric methods, and manipulate marginal distributions. It is designed to enable scientists and practitioners to handle labelled multivariate data typical of computational sciences. For example, given some vertical profiles of atmospheric temperature, we can use Synthia to generate new but statistically similar profiles in just three lines of code (Table 1).

Synthia supports three methods of multivariate data generation through: (i) fPCA, (ii) parametric (Gaussian) copula, and (iii) vine copula models for continuous (all), discrete (vine), and categorical (vine) variables. It has a simple and succinct API to natively handle xarray's labelled arrays and datasets. It uses a pure Python implementation for fPCA and Gaussian copula, and relies on the fast and well tested C++ library vinecopulib through pyvinecopulib's bindings for fast and efficient computation of vines. For more information, please see the website at https://dmey.github.io/synthia.

Table 1. Example application of Gaussian and fPCA classes in Synthia. These are used to generate random profiles of atmospheric temperature similar to those included in the source data. The xarray dataset structure is maintained and returned by Synthia.

Source	Synthetic with Gaussian Copula	Synthetic with fPCA
`ds = syn.util.load_dataset()`	`g = syn.CopulaDataGenerator()`	`g = syn.fPCADataGenerator()`
	`g.fit(ds, syn.GaussianCopula())`	`g.fit(ds)`
	`g.generate(n_samples=500)`	`g.generate(n_samples=500)`

Documentation

For installation instructions, getting started guides and tutorials, background information, and API reference summaries, please see the website.

How to cite

If you are using Synthia, please cite the following two papers using their respective Digital Object Identifiers (DOIs). Citations may be generated automatically using Crosscite's DOI Citation Formatter or from the BibTeX entries below.

Synthia Software	Software Application
DOI: 10.21105/joss.02863	DOI: 10.5194/gmd-14-5205-2021

@article{Meyer_and_Nagler_2021,
  doi = {10.21105/joss.02863},
  url = {https://doi.org/10.21105/joss.02863},
  year = {2021},
  publisher = {The Open Journal},
  volume = {6},
  number = {65},
  pages = {2863},
  author = {David Meyer and Thomas Nagler},
  title = {Synthia: multidimensional synthetic data generation in Python},
  journal = {Journal of Open Source Software}
}

@article{Meyer_and_Nagler_and_Hogan_2021,
  doi = {10.5194/gmd-14-5205-2021},
  url = {https://doi.org/10.5194/gmd-14-5205-2021},
  year = {2021},
  publisher = {Copernicus {GmbH}},
  volume = {14},
  number = {8},
  pages = {5205--5215},
  author = {David Meyer and Thomas Nagler and Robin J. Hogan},
  title = {Copula-based synthetic data augmentation for machine-learning emulators},
  journal = {Geoscientific Model Development}
}

If needed, you may also cite the specific software version with its corresponding Zendo DOI.

Contributing

If you are looking to contribute, please read our Contributors' guide for details.

Development notes

If you would like to know more about specific development guidelines, testing and deployment, please refer to our development notes.

Copyright and license

Acknowledgements

Special thanks to @letmaik for his suggestions and contributions to the project.

Comments

Explain how to run the test suite
Describe the bug There is a test suite, but the documentation does not explain how to run it.

Here is what works for me:

Install pytest.

Clone the source repository.

Run pytest in the root directory of the repository.
opened by khinsen 7
Review: Copula distribution usage and examples

Your package offers support for simulating vine copulas. However, I don't see examples demonstrating how to simulate data from a vine copula given desired conditional dependency requirements.

Is this possible with the current API? If not, how would I use the vine copula generator to achieve this?

Otherwise, can examples show the difference between simulating Gaussian and vine copulas? I only see examples for the Gaussian copula.

opened by mnarayan 5
fPCA documentation
Describe the bug

The documentation page on fPCA says:

PCA can be used to generate synthetic data for the high-dimensional vector $X$. For every instance $X_i$ in the data set, we compute the principal component scores $a_{i, 1}, \dots, a_{i, K}$. Because the principal components $v_1, \dots, v_K$ are orthogonal, the scores are necessarily uncorrelated and we may treat them as independent.

The claim that "because the principal components $v_1, \dots, v_K$ are orthogonal, the scores are necessarily uncorrelated" looks wrong to me. These scores are projections of the $X_i$ onto the elements of an orthonormal basis. That doesn't make them uncorrelated. There are lots of orthonormal bases one can project on, and for most of them the projections are not uncorrelated. You need some property of the distribution of $X$ to derive a zero correlation, for example a Gaussian distribution, for which the PCA basis yields approximately uncorrelated projections.
opened by khinsen 3
Review: Clarify API

It would be helpful to add/explain what the different classes do Data Generators, Parametrizer, Transformers somewhere in the introduction or usage component of the documentation. Explain the different classes and what each is supposed to do. If it is similar to or inspired by well-known API of a different package, please point to it.

I think generators and transformers are obvious but I only sort of understand Parametrizers. It is also confusing in the sense that people might think this has something to do with parametric distributions when you mean it to be something different.

Is this API for Parametrizers inspired by some convention elsewhere? If so it would be helpful to point to that. For instance, the generators are very similar to statsmodel generators.

opened by mnarayan 2
Small error in docs

Hi, just letting you know I noticed a small error in the documentation.

At the bottom of this page https://dmey.github.io/synthia/examples/fpca.html

The error is in line [6] of the code, under "Plot the results".

You have: plot_profiles(ds_true, 'temperature_fl')

But I believe it should be: plot_profiles(ds_synth, 'temperature_fl')

you want to plot results, not the original here.

Cheers & thanks for the cool project!

opened by BigTuna08 1
Review: Comparisons to other common packages

What are other packages people might use to simulate data (e.g. statsmodels comes to mind) and how is this package different? Your package supports generating data for multivariate copula distributions and via fPCA. I understand what this entails but I think this could use further elaboration.

This package supports nonparametric distributions much more than the typical parametric data generators found in common packages and it would be useful to highlight these explicitly.

opened by mnarayan 1
Support categorical data for pyvinecopulib

During fitting, category values are reindexed as integers starting from 0 and transformed to one-hot vectors. The opposite during generation. Any data type works for categories, including strings.

opened by letmaik 0

Add support for categorical data

We can treat categorical data as discrete but first we need to pre-process categorical values by one hot encoding to remove the order. Re API we can change the current version from

# Assuming  an xarray datasets ds with X1 discrete and and X2 categorical 
generator.fit(ds, copula=syn.VineCopula(controls=ctrl), is_discrete={'X1': True, 'X2': False})

to something like

with X3 continuous 
g.fit(ds, copula=syn.VineCopula(controls=ctrl), types={'X1': 'disc', 'X2': 'cat', 'X3': 'cont'})

opened by dmey 0

Add support for handling discrete quantities
Introduces the option to specify and model discrete quantities as follows:

# Assuming an xarray datasets ds with X1 discrete and and X2 continuous generator.fit(ds, copula=syn.VineCopula(controls=ctrl), is_discrete={'X1': True, 'X2': False})

This option is only supported for vine copulas
opened by dmey 0

Releases(1.1.0)

1.1.0(Sep 1, 2021)
Pin pyvinecopulib version to avoid issues between versions.

Add CI tests for Python 3.9 (#17).

Minor doc improvements.

Source code(tar.gz)
Source code(zip)
1.0.0(Apr 19, 2021)
1.0.0

Add JOSS summary paper (#26).

Improve docs and tutorials (#14, #13, #18, ...).

Enable CI on multiple OS and Python versions (#16).

Source code(tar.gz)
Source code(zip)
0.3.0(Nov 12, 2020)
Add support for handling categorical quantities (#10, #13).

Source code(tar.gz)
Source code(zip)
0.2.0(Nov 11, 2020)
Add support for handling discrete quantities (#9).

Add support for setting a seed when generating new samples (#11).

Drop support for Python 3.7.

Source code(tar.gz)
Source code(zip)
0.1.1(Oct 24, 2020)
Fix qrng argument for pyvinecopulib due to vinecopulib/pyvinecopulib#68 and vinecopulib/pyvinecopulib#69.

Source code(tar.gz)
Source code(zip)
0.1.0(Oct 22, 2020)
First public release.

Source code(tar.gz)
Source code(zip)

Owner

GitHub Repository https://dmey.github.io/synthia

ELFXtract is an automated analysis tool used for enumerating ELF binaries

ELFXtract ELFXtract is an automated analysis tool used for enumerating ELF binaries Powered by Radare2 and r2ghidra This is specially developed for PW

49 Nov 28, 2022

First and foremost, we want dbt documentation to retain a DRY principle. Every time we repeat ourselves, we waste our time. Second, we want to understand column level lineage and automate impact analysis.

dbt-osmosis First and foremost, we want dbt documentation to retain a DRY principle. Every time we repeat ourselves, we waste our time. Second, we wan

150 Jan 06, 2023

The micro-framework to create dataframes from functions.

762 Jan 07, 2023

Pizza Orders Data Pipeline Usecase Solved by SQL, Sqoop, HDFS, Hive, Airflow.

PizzaOrders_DataPipeline There is a Tony who is owning a New Pizza shop. He knew that pizza alone was not going to help him get seed funding to expand

4 Jun 05, 2022

A fast, flexible, and performant feature selection package for python.

linselect A fast, flexible, and performant feature selection package for python. Package in a nutshell It's built on stepwise linear regression When p

88 Dec 06, 2022

Gathering data of likes on Tinder within the past 7 days

tinder_likes_data Gathering data of Likes Sent on Tinder within the past 7 days. Versions November 25th, 2021 - Functionality to get the name and age

12 Jan 05, 2023

MotorcycleParts DataAnalysis python

We work with the accounting department of a company that sells motorcycle parts. The company operates three warehouses in a large metropolitan area.

1 Jan 12, 2022

Udacity - Data Analyst Nanodegree - Project 4 - Wrangle and Analyze Data

WeRateDogs Twitter Data from 2015 to 2017 Udacity - Data Analyst Nanodegree - Project 4 - Wrangle and Analyze Data Table of Contents Introduction Proj

1 Jan 12, 2022

Python scripts aim to use a Random Forest machine learning algorithm to predict the water affinity of Metal-Organic Frameworks

The following Python scripts aim to use a Random Forest machine learning algorithm to predict the water affinity of Metal-Organic Frameworks (MOFs). The training set is extracted from the Cambridge S

1 Jan 09, 2022

Synthetic data need to preserve the statistical properties of real data in terms of their individual behavior and (inter-)dependences

Related tags

Overview

Overview

Documentation

How to cite

Contributing

Development notes

Copyright and license

Acknowledgements

Comments

you want to plot results, not the original here.

Releases(1.1.0)

1.1.0(Sep 1, 2021)

1.0.0(Apr 19, 2021)

0.3.0(Nov 12, 2020)

0.2.0(Nov 11, 2020)

0.1.1(Oct 24, 2020)

0.1.0(Oct 22, 2020)

Owner

ELFXtract is an automated analysis tool used for enumerating ELF binaries

First and foremost, we want dbt documentation to retain a DRY principle. Every time we repeat ourselves, we waste our time. Second, we want to understand column level lineage and automate impact analysis.

The micro-framework to create dataframes from functions.

Pizza Orders Data Pipeline Usecase Solved by SQL, Sqoop, HDFS, Hive, Airflow.

A fast, flexible, and performant feature selection package for python.

Gathering data of likes on Tinder within the past 7 days

MotorcycleParts DataAnalysis python

Udacity - Data Analyst Nanodegree - Project 4 - Wrangle and Analyze Data

Python scripts aim to use a Random Forest machine learning algorithm to predict the water affinity of Metal-Organic Frameworks

Python beta calculator that retrieves stock and market data and provides linear regressions.

A notebook to analyze Amazon Recommendation Review Dataset.

PyEmits, a python package for easy manipulation in time-series data.

Automated Exploration Data Analysis on a financial dataset

Pipeline to convert a haploid assembly into diploid

simple way to build the declarative and destributed data pipelines with python

Using approximate bayesian posteriors in deep nets for active learning

Produces a summary CSV report of an Amber Electric customer's energy consumption and cost data.

Extract data from a wide range of Internet sources into a pandas DataFrame.

Data Analytics on Genomes and Genetics

Integrate bus data from a variety of sources (batch processing and real time processing).