Skip to content

Author

Dmitri Kireev

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

DEL-GPT: learning the language of DNA-encoded libraries to design focused screening collections

Focused DNA-encoded libraries (DELs) yield higher hit rates and cleaner selection data than billioncompound collections, and are affordable enough to build for a single target. Designing one means choosing a few hundred synthons from a much larger pool, and in the split-and-pool format that choice is collective and irreversible: the library is the full combinatorial product, so each synthon's value depends on every other one chosen alongside it. Brute force is out of reach – there are on the order of 10189 ways to draw 100 synthons from 3,000. We therefore recast the problem as sequence modeling. Treating a synthon as a token and an efficiency-ordered synthon set as a sentence, a generative pretrained transformer trained by next-token prediction learns context-dependent synthon value and composes new sets one synthon at a time. DEL-GPT was trained on one million ranked synthon sets drawn from NADEL, a validated 58,302-member library, with one model per target for five NAD+-dependent enzymes (PARP1, PARP2, PARP10, PARP12, PARP15). Across all five, DEL-GPT libraries outperformed those built from randomly drawn synthons – a demanding baseline, since every NADEL synthon is an expert-selected NAD+ mimetic – and contained two to three times more compounds that experimental affinity selection retained. Quality falls off once generated sets exceed the size range seen in training, but a model trained on longer sequences generated correspondingly larger high-quality sets – a validated recipe for scaling. The method is chemistryagnostic and applies to any combinatorial library over a shared building-block set.

Akhila Mettu, Naveed Naemi, Raphael Franzini et al. · 0 citations