Synthetic Data: How Generative AI Models Create New Realities

Data & AI
inventx-synthetische-daten-generative-ki-modelle
The responsible handling of sensitive customer data is one of the key challenges in the financial and insurance sector. At the same time, the need for realistic data is growing in order to test systems quickly, continuously improve software, and effectively train AI applications. Synthetic data offers valuable support here. With the latest advances in generative AI, new possibilities are emerging for its creation. InventxLab has taken a closer look at this topic and shares an assessment in this blog post.

What are synthetic data – and why are they relevant? 

Synthetic data are artificially generated datasets that replicate the structure and characteristics of real data, yet do not allow any conclusions to be drawn about real individuals or transactions. They open up new possibilities, especially where data protection, data availability, and quality assurance play a central role. 

The possible applications of synthetic data are broad and cover a wide range of use cases in the banking and insurance industry: 

  1. Testing and quality assurance 

    During development and release phases, testing with real data is not always possible—either because of data protection requirements or because suitable data simply do not exist. Synthetic data enable safe, realistic tests for continuous quality assurance of new software.

  2. Training AI models 

    Training AI and machine-learning algorithms requires large amounts of data. To improve the training of these models, synthetic data can, on the one hand, fill gaps and, on the other hand, generate sufficient data for model training.

  3. Simulation and load testing

    Performance and load tests require realistic stress scenarios. With synthetic data, these can be replicated easily and safely, regardless of the data volume. 

  4. Software development

    Development teams benefit from flexible, data-protection-compliant test data in early project phases. This accelerates development cycles and increases software quality.

  5. Data protection

    Unlike pseudonymisation or anonymisation tools, synthetic data offer an additional layer of security, as they do not allow conclusions to be drawn about real individuals. 

 

Creating new data with generative AI

Synthetic data generation can be carried out using various methods—such as rule-based approaches, simulations, statistical models, or with the help of artificial intelligence. With generative AI, even complex, text-based datasets can be created. The challenge remains complex relationships, e.g. between customer and transaction data or time series, as are typical for the financial industry. Can generative AI handle that? 

The quality of synthetic data generated with generative AI depends heavily on the models used. Depending on the data type, volume, and complexity, different models are suitable: 
 

  • GANs (Generative Adversarial Networks): 

    Two neural networks (generator and discriminator) compete against each other to produce realistic data. Particularly suitable for tabular data, images, and time series. 

  • VAEs (Variational Autoencoders): 

    These models learn a compressed representation of the data and can generate new, similar data points from it. 

  • LLMs (Large Language Models): 

    For unstructured data such as text or complex tables, large language models are increasingly being used, as they are capable of generating realistic and diverse datasets. 

  • TabularAR-GN and other specialised models: 

    For multi-table data or very large, heterogeneous datasets, specialised architectures are used that can represent relationships between tables. 
     

A challenge remains multi-table data, which are needed, for example, for fraud detection in payment transactions or for testing functionalities in the core banking system. This is shown by the InventxLab study conducted as part of a proof of concept on the quality of various providers and models. The selection of suitable models for this is currently still limited. Many providers offer only a single model for multi-table data. In addition, the quality of the models varies greatly. Given the rapid technological development, it is foreseeable, however, that in the near future the selection for these specific models will grow. 

The market for AI-based synthetic data generation is still being established, but already offers a variety of specialised platforms and services with user-friendly operation. Depending on the data type, volume, and complexity, a suitable model is selected and trained on the real data. Afterwards, it can continuously generate new synthetic datasets at the push of a button. 
As part of the provider screening, it is noticeable that primarily specialised providers offer AI-based synthetic data generation. More broadly positioned platforms, by contrast, often still rely on classical methods for generating synthetic data. 
In our view, the potential of AI-supported synthetic data generation remains undisputed and will continue to gain importance as model diversity and quality increase. 

Secure innovation thanks to Swiss infrastructure

A prerequisite for generating synthetic data with AI is a secure, high-performance infrastructure. Inventx therefore relies on GPU-accelerated computing power in its own Swiss data centres. The new AI platform makes it possible to train AI models securely and generate synthetic data in a protected environment, without sensitive information leaving the environment. 

This creates the foundation for Inventx to support its customers in the responsible use of AI and data innovation. 

Find out more in our media release on the GPU service for AI applications.
 

Author

Carla Caspar

Product Manager Data Platform & AI Services

LinkedIn
Foto Carla Caspar

Contact