Skip to content
Open access

Measuring the quality and efficiency of AI-generated codes for financial markets prediction with LSTM

Jul 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 26 references
Medicine

TL;DR

This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini 2.0 Pro, Claude 3.7 Sonnet, and Meta's Llama from a single standardized prompt, on three indices.

Abstract

Generative AI coding assistants are increasingly used to write machine-learning code, yet their ability to produce reliable LSTM implementations for financial prediction remains underexplored. This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini 2.0 Pro, Claude 3.7 Sonnet, and Meta’s Llama from a single standardized prompt, on three indices (Nikkei 225, S&P 500, STOXX Europe 600). Each assistant’s generated script was re-executed over independent runs; accuracy (MAE, MSE, RMSE, R2, execution time) is reported as mean ± standard deviation on the original price scale, complemented by a static code-quality analysis (Pylint, Radon, SonarQube, Pytest, Bandit). The assistants converge on nearly identical LSTM architectures, so performance differences arise mainly from data-handling and code-correctness defects: Meta’s Llama near-zero errors are an artifact of normalized-scale metrics combined with a shuffled train/test split (data leakage), and once corrected its accuracy is among the weakest; Gemini 2.0 Pro, once its predictions are evaluated consistently on the price scale, is among the most accurate assistants. Differences are validated with Diebold–Mariano and Wilcoxon tests. AI-generated forecasting code can be accurate but is not uniformly trustworthy: its generated preprocessing and evaluation code must be audited before use.

Read PDF

Similar papers

Jul 2026

AI Trading: Evaluating Large Language Models for Technical Market Analysis

It is concluded that while LLMs hold genuine promise within AI trading systems, robust deployment requires careful task decomposition, rigorous backtesting protocols, and domain-aware fine-tuning strategies.

Geofrey Ntale · 0 citations
Open access Aug 2026

Deep Learning vs. Traditional Machine Learning for Software Defect Prediction: A Meta-Analysis

: Deep learning (DL) is now routine in software defect prediction (SDP), yet how much it improves on traditional machine learning (ML), how stable that improvement is, and what governs it remain contested. We synthesized 45 empirical studies published between 2015 and 2024, comprising 1540 performance estimates, using...

Wei-Xiang Gan, Jia-Lin Liu, Meng-Fei Xiao et al. · 0 citations
Open access Aug 2026

Pengembangan Model Klasifikasi Code Smells Pada Backend Python Menggunakan Algoritma Random Forest (Studi Kasus Proyek Open Source Github)

Deadline pressure in software development often drives coding shortcuts, leading to internal quality degradation known as code smells. These structural anomalies contribute to technical debt accumulation and complicate system maintenance over time. This study develops an automated classification model to detect code sm...

Unknown authors · 0 citations
Open access Sep 2026

Predicting Startup Outcomes Using Explainable Machine Learning and Y Combinator-Inspired Feature Engineering

Early-stage startup success is notoriously difficult to predict, yet the stakes for getting it right - whether for investors, accelerators, or founders themselves - are substantial. Most existing machine learning approaches lean heavily on generic features from Crunchbase or PitchBook, and in doing so tend to miss a ca...

Siddharth Gupta, Pratham Namdev, Shubham Nagar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.