Extracting structured information from massive unstructured texts is a key task for intelligent analysis and knowledge construction, but faces two challenges in low-annotation-resource scenarios: the high cost of obtaining high-quality labeled data and the prevalence of domain-specific terminology and complex semantic structures. Existing methods still face several challenges, including static knowledge transfer, insufficient modeling of learner capability differences, and limited consideration of sample difficulty during teaching strategies. This paper proposes the Dynamic Loop Teaching Method (DLTM), aiming to enhance the information extraction capability of small language models under extreme few-shot conditions. DLTM constructs three identical small models with differentiated initial capacities. Through multi-round dynamic mutual teaching, they adaptively assume asymmetric roles: the relatively stronger model acts as the Optimizer, providing pseudo-supervision signals; the intermediate model acts as the Questioner, generating targeted samples; and the weaker model acts as the Respondent, being fine-tuned. Key components include dynamic role assignment, validation set expansion with rollback protection, and adaptive difficulty curriculum learning.Experiments are conducted on a publicly available Chinese event extraction dataset. The results show that DLTM enhances the ability of small models to extract complex semantic structures and achieves competitive performance compared with representative baselines. In particular, DLTM achieves 17%–25% relative improvement on coreference resolution under corresponding settings compared with ODIE and ADELIE. Further analysis indicates that while DLTM improves semantic-dependent extraction tasks, knowledge-intensive fields remain challenging under extremely sparse supervision.
Wei-Hang Du, Yao-Hong Zhang, Da-Yu Zhang et al.· 2026 12th International Conf...· 0 citations
A key problem in language-guided UAV target search is how to transform a language-referred target in the current observation into an executable spatial goal. Existing methods either predict actions directly or introduce relatively heavy mapping, memory, or planning modules, making the intermediate link between semantic grounding and spatial execution difficult to examine in isolation. In this paper, we present a lightweight closed-loop framework for language-guided UAV target search and reaching. Given a natural-language instruction, an RGB image, a depth map, and the UAV pose, the system first localizes a 2D target with a vision-language model, then recovers a 3D search goal in the world coordinate system using depth cues and camera geometry, and finally executes point-to-point flight toward the recovered goal. Rather than addressing obstacle avoidance, global mapping, cooperative coverage, or complex trajectory optimization, we focus on validating whether semantic target grounding, explicit 3D search-goal recovery, and flight execution can form an effective perception-to-execution loop. Preliminary AirSim results show that observation-consistent 3D search-goal recovery yields more stable target-search execution than both an image-plane heuristic baseline and a fixed-depth recovery baseline. These results suggest that explicit 3D search goals provide a practical and interpretable bridge between semantic grounding and spatial execution.
Jun-Song Zhang, Yao-Hong Zhang, Rui Guan et al.· 2026 12th International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.