This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini 2.0 Pro, Claude 3.7 Sonnet, and Meta's Llama from a single standardized prompt, on three indices.
Abstract
Generative AI coding assistants are increasingly used to write machine-learning code, yet their ability to produce reliable LSTM implementations for financial prediction remains underexplored. This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini 2.0 Pro, Claude 3.7 Sonnet, and Meta’s Llama from a single standardized prompt, on three indices (Nikkei 225, S&P 500, STOXX Europe 600). Each assistant’s generated script was re-executed over independent runs; accuracy (MAE, MSE, RMSE, R2, execution time) is reported as mean ± standard deviation on the original price scale, complemented by a static code-quality analysis (Pylint, Radon, SonarQube, Pytest, Bandit). The assistants converge on nearly identical LSTM architectures, so performance differences arise mainly from data-handling and code-correctness defects: Meta’s Llama near-zero errors are an artifact of normalized-scale metrics combined with a shuffled train/test split (data leakage), and once corrected its accuracy is among the weakest; Gemini 2.0 Pro, once its predictions are evaluated consistently on the price scale, is among the most accurate assistants. Differences are validated with Diebold–Mariano and Wilcoxon tests. AI-generated forecasting code can be accurate but is not uniformly trustworthy: its generated preprocessing and evaluation code must be audited before use.
It is concluded that while LLMs hold genuine promise within AI trading systems, robust deployment requires careful task decomposition, rigorous backtesting protocols, and domain-aware fine-tuning strategies.
: Deep learning (DL) is now routine in software defect prediction (SDP), yet how much it improves on traditional machine learning (ML), how stable that improvement is, and what governs it remain contested. We synthesized 45 empirical studies published between 2015 and 2024, comprising 1540 performance estimates, using...
Wei-Xiang Gan, Jia-Lin Liu, Meng-Fei Xiao et al.· Engineering and Computing In...· 0 citations
Deadline pressure in software development often drives coding shortcuts, leading to internal quality degradation known as code smells. These structural anomalies contribute to technical debt accumulation and complicate system maintenance over time. This study develops an automated classification model to detect code sm...
Treating supervision format as a first-class hyperparameter for multi-task reasoning SFT in large language models—at least in this benchmark-and-model setting—rather than a mere rendering detail is supported.
Nhat Thanh Vu, M. Rashid, Fariza Sabrina· Electronics· 0 citations
Early-stage startup success is notoriously difficult to predict, yet the stakes for getting it right - whether for investors, accelerators, or founders themselves - are substantial. Most existing machine learning approaches lean heavily on generic features from Crunchbase or PitchBook, and in doing so tend to miss a ca...
Siddharth Gupta, Pratham Namdev, Shubham Nagar et al.· Cureus Journal of Computer S...· 0 citations
From the ways agents exploited their harness--reading sibling runs through shared git state, leaving notes to"future runs"in persistent memory--the authors distill five design rules for evaluating autonomous agents.
N. Askarbekuly, Mohamad Al Mdfaa, Ahmed Helaly et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.