How do you feel about the legality and morality of “borrowing” copyrighted data for machine learning purposes?
This has become somewhat of a hot topic these days, largely due to mass exploitation of human cultural output by companies which train and host large language models. As such, I’m worried of writing something with nuance because of how heated the conversation around it has become, particularly due to concerns over the impact on labor and employment1. However, I’m not here to discuss generative AI or LLMs, not directly at least. Rather, I’d like to try and shine a different light on the broader issues of data mining, through an example rooted in a popular technology which has enabled an explosion of creative approaches in the landscape of online streaming.
One day, some years ago, I decided to look at the data used to train OpenSeeFace. OpenSeeFace is the most popular open source face tracking solution for virtual YouTubers. It is supported by both open source and commercial model rendering tools; in particular, VTube Studio bundles it as an option for webcam tracking.