Principles for building a product based on generative neural network models
Abstract
The availability of generative models through cloud APIs has lowered the barrier to building products around them, but it has not removed the underlying engineering problem: the mere ability to call a generative model does not guarantee that a workable, scalable, and degradation-resistant product can be built around it. The article systematizes the principles of building such a product – from selecting and combining generative models and designing a multiplatform delivery architecture to the methodology of testing technical hypotheses and the metrics used to evaluate generation quality. Using the case of a multiplatform service that generates personalized visual content, the article shows how the sequential testing of technical hypotheses – moving from simple face-swap to controllable generation with a fine-tuned diffusion model, automatic enrichment of user prompts, an image-to-image generation feature, compute optimization, and a cross-platform architecture – enables the transition from a prototype to a product that remains stable under a multi-fold increase in load. The article separately examines the principles of quality assurance for generative output, technical infrastructure scaling, and the mitigation of risks specific to generative AI products: dependency on a single model provider, quality degradation under load, and the vulnerability of fixed pricing to fluctuations in inference cost.