This work retrain eight open models from their base weights on a shared pool of each other's text for five generations, varying the part written by one model, Phi-2, from an equal share to 90%.
What kind of data does a model need in order to learn? Coreset selection makes this question concrete: under a budget, keep the samples most useful for training. Easy-first and geometric coverage criteria can win in different budget regimes, separated by a crossover boundary. We ask whether this boundary is fixed by th...
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmar...
AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another -- but almost always with the market split evenly, while real generative AI is an oligopoly. Concentration...
Yang-Ze Liu, Zhong-Yi Han· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.