These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures, and that CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi et al.· 0 citations
The first comprehensive benchmarking framework specifically designed to accommodate inter-dataset heterogeneity is presented, finding that well-designed small datasets can match or even surpass the performance of larger benchmarks, suggesting that different metrics are applicable to different datasets/testing scenarios.
Yingjuan Cheng, Qing Ye, Linlong Jiang et al.· Journal of Cheminformatics· 0 citations
A novel committor learning framework grounded in the AlphaFold 3 paradigm is proposed that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.
Jintu Zhang, Zichang Jin, Huifeng Zhao et al.· 0 citations
A novel approach for distinguishing individuals with Autism Spectrum Disorder (ASD) using Intuitionistic Fuzzy Set (IFS) theory and Multi-Scale Enhanced Graph Convolutional Networks (MSE-GCNs), which represents a substantial improvement over existing models for ASD and potentially for other neurological disorders.
S. Rajaprakash, C. Basha, K. Manivanan et al.· Discover Artificial Intellig...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.
Mingze Yin, Yiheng Zhu, Jialu Wu et al.· Proceedings of the 32nd ACM...· 0 citations
This work presents a multi-mode energy harvesting-assisted edge computing architecture, integrated with a joint optimization of energy consumption and communication behaviour, aimed at enhancing the sustainability, reliability and autonomy of operation in an industrial IoT context.
Dr. Deepa, M. Mehfooza, Padmavathy Thiruppathi Raj· Microsystem Technologies· 0 citations
It is argued that AI-ML integration improves productivity in agriculture in terms of crop yield prediction, disease prediction and optimization of resources, amongst others, and a comprehensive strategy for future work in designing sustainable agrifood systems is proposed.
Anita Veerappa Karkikatti, R. H. Goudar, Vijayalaxmi N. Rathod et al.· Discover Artificial Intellig...· 0 citations
Scalability analysis demonstrates that the proposed Post-Quantum Probabilistic Hidden-State Deep Learning framework, evaluated with run on IoT networks with over 1000 nodes, exhibits significant performance.
T. G. Keshavamurthy, S. Guruprasad, K. Hareesh et al.· Discover Artificial Intellig...· 0 citations
This study demonstrates the efficacy of the synergy between federated learning and edge computing in IoT security contexts, providing a scalable and privacy-centric solution for anomaly detection across large-scale distributed devices.
Quan Liu, Yuanyuan Feng· Discover Artificial Intellig...· 0 citations
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.
Desiree Cho, Cameron Tice, Bernie Hogan et al.· 0 citations
The first $poly(\Delta,\log n)-round algorithm for $(\Delta + 1)$-edge coloring in the CONGEST model is presented and the $n$-dependency of its runtime, $\tilde{O}(\log^5 n)$, matches the best published dependency in the LOCAL model.
Sebastian Brandt, Ananth Narayanan, Alexandre Nolin· 0 citations
The Ramsey number $R(m,n)$ is the smallest order at which every red-blue edge coloring of a complete graph must contain a blue clique (a complete subgraph) of size $m$ or a red clique of size $n$. Determining these numbers exactly is extremely hard, and even certifying a lower bound requires exhibiting an explicit coloring that avoids both cliques. We develop an integer programming framework for certifying such lower bounds, restricting the search to circulant graphs, whose rotational symmetry lets us reformulate the problem in a projected distance space, reducing the number of binary variables from quadratic to linear in the graph order. We strengthen this projected model through coefficient reduction and solve it with a branch-and-cut algorithm whose separation routine exploits the common neighborhood structure of circulant graphs, combining heuristic and exact maximum-clique algorithms. In an extensive computational campaign on circulant graphs with up to 410 vertices, we improve the best lower bounds previously obtained by other methods by up to 11 points for 25 values of $R(3,n)$ with $24\le n\le49$ and $n\neq27$, each backed by an explicit graph certificate that can be independently verified with a stand-alone exact clique solver. To the best of our knowledge, our method also provides the first reproducible optimization-based procedure for certifying circulant Ramsey numbers $R_C(m,n)$, which we use to establish eight new values of $R_C(3,n)$ with $13\le n\le20$. Our framework, graph certificates, and stand-alone checker are provided as supplementary material to support independent verification and reuse.
Stefano Coniglio, Fabio Furini, I. Ljubić et al.· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.