Skip to content
Open access

MOCA: A Hierarchical Semantic-Enhanced Code Edit Framework for Multilingual Code Co-Evolution

Abstract

Multilingual code co-evolution aims to propagate code edits across programming languages while preserving functional consistency, yet the task remains challenging in real-world repositories. Our motivating examples reveal two key challenges: (1) Existing models fail to infer the rationale behind source edits and may omit or misidentify the corresponding edits in the target code; and (2) Models do not capture repository-level semantics such as API usages and cross-file dependencies required for correct adaptations. To address these limitations, we introduce MOCA, a hierarchical semantic-enhanced framework that operationalizes two complementary strategies: (1) edit-wise intent interpretation, which explicates the intent of each source edit and guides its cross-lingual mapping, and (2) project-wise context integration, which injects repository-level semantics to ensure that propagated edits adhere to project constraints. MOCA implements these strategies through four coordinated steps: a Summarizer that provides multi-perspective explanations of each source edit, a Locator that identifies corresponding edit regions in the target code, a Retriever that supplies and filters repository context, and a Modifier that integrates all reasoning signals to synthesize the final edits. Experiments on eight real-world Java–C# projects demonstrate that MOCA achieves substantial improvements under both unseen-project ( \(S_{proj}\) ) and time-split seen-project ( \(S_{time}\) ) settings, reaching 73.06%/71.84% exact match for C#→Java/Java→C# under \(S_{proj}\) and 85.99%/79.36% under \(S_{time}\) . Compared with the state-of-the-art fine-tuning-based method Codeditor, MOCA improves exact match by 11.69–32.04%; compared with prompt-engineering-based baselines, it further yields 5.10–6.54% under \(S_{proj}\) and 5.63–5.81% under \(S_{time}\) . Our ablation studies confirm the contribution of all designed agents, while generalization experiments show that MOCA consistently outperforms each model's strongest prompt baseline by 3.56–15.63%. Qualitative analyses further indicate that MOCA handles a broader range of multi-hunk and semantically complex edits. Overall, MOCA provides a practical paradigm for multilingual code co-evolution by jointly strengthening edit-wise semantics and project-wise grounding.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.