Researchers at Tsinghua University and ByteDance have posted a preprint, a paper not yet peer reviewed, describing Timer-M1, a model for predicting future values in data that changes over time. They report that it ranks first on two forecasting benchmarks. The paper itself says one of those wins is a narrow margin.
What a time-series foundation model is
A time series is a run of measurements over time, such as sales or power demand. A time-series foundation model is trained once on a very large collection of such data, so it can forecast new data without being retrained, which the paper calls “zero-shot” forecasting [1].
What the authors report
Reported (authors’ preprint, arXiv, submitted 8 October 2026; not peer reviewed; self-reported benchmarks). The abstract says: “Across three large-scale forecasting benchmarks, Timer-M1 ranks first on both FEV and TIME and second on GIFT-Eval among most recent time series foundation models.” [1] On FEV, it reports “MASE skill 0.3771 and SQL skill 0.4887 over the 100 matched configurations”, with a “compact 121M-parameter backbone” (121 million adjustable settings) [1]. The skill scores measure error against a simple baseline forecast, and higher is better. A configuration is one test setup in the benchmark.
The authors’ own caveats
- TIME. The margin over Google Research’s TimesFM-3 is small. The paper says: “We interpret this narrow margin as competitive performance with TimesFM-3 rather than statistically established superiority.” [1]
- GIFT-Eval. Timer-M1 ranks second, “while TimesFM-3 retains the lowest aggregate errors”. [1]
- Training data. “Training-data exposure is not matched across the compared methods.” [1] The models may have been trained on different data, so the scores may not compare like with like.
- The rankings. The paper says the ranks “refer to the displayed methods rather than a claim of official zero-shot leaderboard placement”. [1]
What this does not show
- Peer review or replication. No outside group has checked these results.
- A clear win over TimesFM-3. On FEV the paper reports gaps of 0.47% and 0.42% in relative error; on GIFT-Eval TimesFM-3 is ahead.
- Performance on your data. Benchmarks are standard test sets, not any one business’s records.
The Bottom Line
The authors report that Timer-M1 tops two of three forecasting benchmarks, and their own paper cautions that the lead over TimesFM-3 is narrow, that TimesFM-3 is ahead on the third, and that training data differ across the models compared. It is a claim to be tested, not a settled ranking.
Sources
- Haoran Zhang, Haixuan Liu, Xingjian Su, Yong Liu, Zhi Chen, Yuxuan Wang, Jianmin Wang and Mingsheng Long, “Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives”, arXiv:2610.11734, submitted 8 October 2026 (preprint; not peer reviewed; self-reported results; all quotes and figures from the abstract and paper). https://arxiv.org/abs/2610.11734

