Current research papers on synthetic examples for training LLMs are often just ideas and lack convincing demonstrations at frontier LLM scale.
I am only at a stage of reading abstracts because a lot of these papers are just ideas. Someone has to make it work on a frontier LLM lab scale in full generality because when you see these papers, they pop up, and it's just a bit noisy. They're cool ideas, but I haven't seen anyone convincingly show that this is possible.