How We Were Inspired
We looked at Kandinsky Lab — encoding Russian culture inside the dataset matrix. LAION and alike datasets have leaned on scale and diversity without taking into account factors that actually matter.
What is the model being trained on? Does it reflect the culture it is being deployed on? How prominent are non-cultural, religious, and sentimental values beyond the expansion of porn and explicit content?
We started Shadow to tackle all these questions and incorporate a swap-in, swap-out mechanism for datasets — where simply taking the dataset should be considered enough for training a large-scale model.
What Parts of the Internet We Open-Source
Russian and Eastern European Internet. We provide data that is already public, plus metadata.
Unlike LAION-BVD and similar datasets where data is locally downloaded and shared with researchers, we open-source critical infrastructure written by our team and push the burden of training — and doing anything with the data — onto the researchers. Based on exemptions and nature, each individual case is different.
What We Plan to Release
- Sleeping-RVD Sleeping-Russian Video Dataset
- Accent1 Large-scale synthetic accent generator and dataset
Whilst these two are the biggest releases, we will release more data and individual projects and code under Project Shadow.
How We Collected the Data
We never used any credentials or logins to access this data. What is being shown is all public internet product.
Do We Release Code?
Actually, no. We feel a responsibility where folks do not abuse our methods for doing malicious things. Research is fine, and we are releasing data. Anything else is off the charts.
Thanks
We thank all our friends and community for helping us.