OpenAI has released 'LifeSciBench,' a benchmark test that measures how useful AI can be to scientists.

OpenAI announced its AI benchmark test, ' LifeSciBench, ' on June 17, 2026. LifeSciBench is a benchmark test that can measure 'how useful AI is for life science researchers,' and is said to be able to provide evaluations that are more in line with actual operations compared to conventional scientific tests.
Introducing LifeSciBench | OpenAI
While several benchmark tests exist to measure AI's performance on science-related tasks, traditional tests often have issues such as 'targeting knowledge in a narrow domain' and 'being in a question-and-answer format with clear correct answers,' which have failed to accurately reflect actual capabilities in the real world.
Therefore, OpenAI classified the tasks that scientists handle on a daily basis into seven categories: 'handling scientific evidence,' 'analysis,' 'design and optimization,' 'scientific consideration,' 'verification and operation,' 'linking scientific findings to clinical decision-making,' and 'scientific communication.' They then collaborated with 173 scientists involved in biotechnology and drug discovery to create tasks. Each task is structured in the form of 'a scientist requesting assistance from a knowledgeable collaborator,' and the AI needs to generate answers in a free-response format while reviewing the content of relevant materials.
LifeSciBench provides the AI with a total of 750 tasks. The AI is given 1062 attachments, including figures, tables, and chemical structure files, and 53% of the tasks are designed to refer to at least one of these attachments.

The AI-generated answers are evaluated against a variety of criteria, such as whether they reach the appropriate level of detail expected by scientists, whether the reasoning is correct, and whether the formatting is correct. Points are added if each criterion is met. This allows LifeSciBench to measure 'how useful AI actually is to scientists.'

The graph below shows the LifeSciBench scores of several AI models, listed in descending order of score as 'GPT-Rosalind,' 'GPT-5.5,' 'Gemini 3.1 Pro,' 'GPT-5.4,' and 'Grok 4.3.' Incidentally, GPT-Rosalind is OpenAI's science-focused AI model, and at the time of writing this article, it is based on GPT-5.5.

The following graph compares the performance of GPT-Rosalind and GPT-5.5. GPT-Rosalind achieved higher scores in all seven tasks.

On the same day as the LifeSciBench announcement, OpenAI released a report stating that 'GPT-5.4 was able to assist in drug discovery research,' highlighting how its proprietary AI is contributing to the advancement of science and technology.
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry | OpenAI
https://openai.com/index/ai-chemist-improves-reaction/

Related Posts:







