# Synthetic Data for ML (Jina AI, Florian Hönicke)

- Channel: [ZuBerlin](https://streameth.org/zuberlin)
- Date: 2024-06-19
- Duration: 32:03
- Topics: computer science, AI, synthetic data, machine learning, data generation
- Watch: https://streameth.org/watch/6672eacb07f92b086c3e1523
- Download: https://vod-cdn.lp-playback.studio/raw/jxf4iblf6wlsyor6526t4tcmtmqa/catalyst-vod-com/hls/56839nye2huwdcrq/1080p0.mp4

## Description

Clip The video features a discussion on synthetic data and its growing importance in training AI models, with predictions that 60% of training data will be synthetic by the end of the year. The speaker emphasizes the benefits of using synthetic data, such as domain specificity, cost reduction, bias control, and consistent labeling. They explore different research papers on generating question-answer pairs for training AI models and discuss various methods and their effectiveness. The talk also covers the limitations of language models in generating unique and domain-specific queries, the superiority of human-generated data in specific domains, and different approaches to improve data generation. The speaker suggests starting with synthetic data generation and shares insights into adversarial example generation as well as the potential of decision-making algorithms. Questions from the audience address concerns about bias amplification, data quality trade-offs, and the efficiency of synthetic versus human-generated data.
