Just another WordPress site - Ruhr-Universität Bochum
From Word2Vec to Transformers: text‐derived composition embeddings for filtering combinatorial electrocatalysts
Compositionally complex solid solution electrocatalysts span vast composition spaces, and even one materials system can contain more candidate compositions than can be measured exhaustively. Here, we compare how five text-derived composition representations influence the outcome of a common label-free filtering task. A focused-corpus Word2Vec model is compared with MatSciBERT and Qwen embeddings, with multicomponent compositions encoded either by concentration-weighted element vectors or by full-composition prompts. Similarities to the concepts conductivity and dielectric define a 2D semantic space in which the same symmetric Pareto-front selection rule is applied to every representation. Across 14 combinatorial HER, ORR, and OER materials libraries, the representations produce distinct trade-offs between candidate reduction and best-current retention. MatSciBERT\_Full gives the smallest subsets on average, whereas Word2Vec gives the smallest mean best-current deviation while retaining comparably compact subsets. These differences show that representation construction changes the behavior of the same downstream filter.