Tsinghua University researchers introduce FuelProp-LM, a novel framework combining instruction tuning and dynamic in-context learning to predict multiple fuel properties directly from molecular SMILES strings.
Predicting the physicochemical properties of fuels is a cornerstone of clean energy transition and high-efficiency engine design. However, many established Quantitative Structure-Property Relationship (QSPR) workflows involve property-specific model development and molecular descriptor selection. Incorporating now measurements may also require model updates or retraining.
In a new study published in ENGINEERING Energy , researchers from the Department of Energy and Power Engineering and the Center for Combustion Energy at Tsinghua University demonstrate how general-purpose open-source Large Language Models (LLMs) can be adapted for fuel property prediction through instruction tuning and in-context learning. The research team developed FuelProp-LM , a unified AI framework capable of predicting multiple complex fuel properties simultaneously using only a molecule’s Simplified Molecular-Input Line-Entry System (SMILES) string as input.
Exploring the Potential of Smaller Language Models To construct FuelProp-LM, the researchers fine-tuned four compact open-source language models with fewer than 10 billion parameters (including Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Phi-3.5-mini-instruct, and DeepSeek-R1-Distill-Qwen-7B) on 100,000 data entries sourced from the PubChem database. Notably, no task-specific fuel property data was used during the instruction fine-tuning stage , allowing the model to learn transferable structural reasoning capabilities.
To enhance contextual inference during prediction, the team integrated a dynamic retrieval algorithm using 166-bit MACCS molecular fingerprints and Tanimoto similarity search via the Hierarchical Navigable Small World (HNSW) algorithm. This approach automatically retrieves structurally similar reference molecules from updated property databases to form informative in-context prompts.
Across comprehensive evaluations on 10 critical fuel property datasets—spanning ignitability, sooting tendency, volatility, and thermodynamics—the fine-tuned 7-billion-parameter FuelProp-LM consistently outperformed the 685-billion-parameter industry-leading DeepSeek-V3.2 baseline . FuelProp-LM also compared favorably with several conventional machine learning baselines.
Key Research Highlights
Gaining Insights into the Model’s “Black Box” To explain why instruction tuning so dramatically improves QSPR accuracy, the researchers conducted an in-depth attention layer analysis. The results demonstrate that fine-tuning fundamentally reallocates internal attention patterns in the intermediate Transformer layers:
"FuelProp-LM bypasses the tedious process of manual descriptor construction and task-specific model retraining," said Professor Bin Yang, corresponding author of the study. "By offering high accuracy, data scalability, and deployment ease, FuelProp-LM provides an accessible and flexible tool for future AI-assisted sustainable fuel design."
About the Research
Journal: ENGINEERING Energy
Read the full article for free: https://rdcu.be/4Xxx7s2Z6Lta
Cite this article: Tao, C., Liu, C., Li, C., & Yang, B. Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning. ENGINEERING Energy , 2026, 20(5): 10760. https://doi.org/10.1007/s11708-026-1076-y
ENGINEERING Energy
News article
Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning
10-Sep-2026