With the problem clearly defined, the next stage is actually assembling a dataset — this note covers where data comes from in practice, and the licensing and ethical considerations that genuinely matter for a real project.
Common Data Sources
| Source | Consideration |
|---|---|
| Public datasets (Kaggle, academic benchmarks, government open data) | Fast to start with, but check licensing terms before any commercial use |
| Web scraping | Can provide large volumes of data, but raises real legal and ethical questions (terms of service, copyright, personal data) that must be addressed explicitly, not assumed away |
| Internal/proprietary data | Often the most task-relevant, but frequently requires careful privacy and compliance handling |
| Synthetic data generation | Useful when real data is scarce or sensitive, but synthetic data's realism and coverage of edge cases must be validated, not assumed |
| Human annotation | Often necessary for supervised tasks lacking existing labels — costly, but sometimes unavoidable |
Rough Data Volume Guidance
There's no universal answer, but as a rough starting intuition: simple classification tasks with a handful of well-separated classes might work reasonably with a few thousand examples per class; complex tasks (fine-grained classification, generative modeling) typically benefit from orders of magnitude more. Transfer learning (the Transfer Learning category) can substantially reduce these requirements by leveraging a model already pretrained on a much larger, related dataset.
Licensing and Ethical Considerations
- Check the license terms of any dataset before commercial use — some public datasets are restricted to research or non-commercial purposes only.
- Consider whether the data contains personally identifiable information, and what privacy/regulatory obligations apply (varies significantly by jurisdiction and data type).
- Consider whether the collection process itself could introduce bias that later shows up as unfair or harmful model behavior — data collection choices directly shape what the model can and will learn.
Common Mistakes
- Using a scraped or public dataset commercially without verifying its actual licensing terms — this can create real legal exposure discovered too late in a project.
- Collecting data without considering whether its composition reflects the real-world population the model will actually be deployed on — a mismatch here (a form of distribution shift, see Domain Adaptation) can silently undermine real-world performance despite strong offline metrics.
Interview Relevance
Q: "What considerations go beyond just 'how much data do I have' when assembling a training dataset?" Licensing terms (especially for commercial use), privacy and regulatory obligations for any personal data, and whether the data's composition genuinely reflects the population the model will be deployed against — a technically large dataset that's poorly matched to real deployment conditions can produce a model that looks strong offline but performs poorly in practice.
Practice Question
Why might a dataset scraped entirely from one specific demographic or geographic region create problems for a model intended to be deployed globally?