The AI Chemist is a column covering what artificial intelligence technologies can do now, what they could do in the future, and what they shouldn’t tackle—all written by expert contributors.
Anyone working with modern artificial intelligence systems will appreciate the extraordinary pace of progress achieved over the last 2 years. In 2024, the frontier models of the time (such as OpenAI’s GPT-4 series) barely reached an IQ equivalent of 100. Today, state-of-the-art systems (including Google 3.x, Anthropic 5.x, and OpenAI 5.x) exhibit IQ equivalents north of 130, according to Tracking AI. Furthermore, explicit reasoning chains (“train of thought”) are now routinely accessible, as well as “agentic” mode (a clever marketing term for AI-powered autonomous bots), which allows for a high degree of task automation.
While AI’s scientific intelligence remains a work in progress, we chemists, along with our colleagues, are likely to benefit most from using AI in specific ways—and in fact, from using different types of AI for different purposes.
The integration of AI has substantially improved my research productivity, albeit not quite in the manner its creators originally envisioned. Large language models (LLMs) have proved exceptionally adept at managing the dense administrative paperwork that modern scientific bureaucracy excels at generating. Before adopting AI, I routinely spent 20% of my work hours stuck in bureaucratic tasks; that burden has now been reduced to 10%, if not less. Many routine paperwork workflows can now be 100% automated.
As a cautionary note, errors remain a feature, not a bug, of all generative AI architectures. This is a substantial concern in society in general and of course carries over to science. Because these systems are inherently creative entities, they can be confidently incorrect in their assertions. Hallucinations and errors can be a serious issue —to the point that researcher John Trant noted in a previous column that he was “nervous enough not to use an LLM.” While the situation has improved dramatically with the introduction of automated cross-checking and verification algorithms, errors cannot be entirely suppressed in principle. Although some errors are easy to catch (such as fabricated citations), detecting more subtle logical flaws requires exhaustive, iterative back-and-forth validation.
When it comes to the core chemical research itself, progress has been more nuanced. Initial attempts to solve complex chemical problems using LLMs were disappointing, largely because of misplaced expectations and flawed project execution. Far better outcomes were later achieved by carefully selecting specific AI architectures for targeted tasks and recalibrating the scope of expected results. Two areas where LLMs have proved quite valuable are in distilling manual chemical preparation procedures into high-throughput robotic synthesis protocols and in identifying potential safety issues with biologically active and energetic compounds.
A major initial misconception was treating all AI systems as fundamentally homogeneous. In reality, two distinct paradigms have emerged under the umbrella acronym of AI: LLMs and large world models (LWMs).
LLMs are built and trained on the corpus of written human knowledge, and they align closely with human intelligence. They excel at writing, classification, and coding. This has proved valuable in the field of technomimetics, where conceptual molecular design is straightforward and shares clear analogies with existing engineering solutions. But the performance of LLMs in less structured, practical chemical tasks, such as optimizing experimental drug development or deploying molecular devices for biological applications, fell short of expectations.
LWMs are trained on raw experimental data and are entirely unconstrained by the boundaries of written human knowledge. In essence, they represent alternative information models of molecular space that possess little to no correlation with human conceptual frameworks. Specialized LWMs, such as AlphaFold descendants, have substantially accelerated our prostate cancer chemotherapy efforts, particularly in hit-to-lead optimization. Remarkably, these systems have also identified valid protein targets for multitarget biologically active molecules that were simply beyond human predictive capacity or intuition.
Despite their utility, LWMs suffer from one fundamental flaw: they are essentially forms of “alien intelligence” and are conceptually incompatible with human reasoning. I have attempted to use LLMs as translators or interpreters for LWM outputs, but those efforts met with mixed results. The information models that LWMs construct are different from our own, and not merely in terms of complexity. For example, certain neural engines treat molecules as information objects residing in a multidimensional vector molecular space that far exceeds the traditional 3D convention. They map intermolecular interaction parameters that cannot be easily categorized within standard human approximations. It remains unknown whether these parameters represent internal mathematical abstractions unique to multidimensional molecular spaces or reveal fundamental molecular features missing from traditional human science.
To bridge this gap, my colleagues and I eventually settled on developing an extended suite of small world models (SWMs), which typically use fewer than 100 parameters. While SWMs lack the sweeping generality of LWMs, their outputs can be partially translated into the realm of human chemistry using LLMs. If developed in sufficient numbers, these SWMs could provide comprehensive and interpretable coverage of the relevant molecular space.
In addition to refining our approach, I’m also excited to see what people who are new to the field bring to both AI and chemistry. I recently served as a judge for AI-driven research projects conducted by high school students and came away deeply impressed by the sophistication and results of their work. The incoming generation will be the first “AI natives,” viewing these systems as a natural extension of human intellect. Notably, these students exhibit no psychological or institutional barriers to adopting AI for day-to-day research tasks. While their relative lack of professional experience will naturally necessitate more rigorous quality control, this is an entirely manageable challenge. Additionally, LLMs are excellent at translation. There’s a possible future in which a potential researcher’s level of fluency in English or any other language is not a barrier to collaboration and communication with a colleague.
In taking a bird’s-eye view of this shifting landscape, I choose to remain cautiously optimistic about the future of the hybrid human-AI model for the chemical sciences. The challenges are formidable, but the opportunities are unprecedented. In this regard, our present epoch is not fundamentally different from the previous industrial revolutions that reshaped human society. AI is now an indelible component of the scientific discovery process; whether an individual researcher chooses to embrace it or reject it is a personal choice.
Credit:
Courtesy of Andrei Gakh
Andrei Gakh is currently a research coordinator at the Department of Energy–sponsored Discovery Chemistry Project, after serving for 2 decades as a senior research scientist at Oak Ridge National Laboratory. His interests include molecular devices, artificial intelligence, and fluorine chemistry.
Views expressed are those of the author and not necessarily those of C&EN or ACS.