Most Viewed

  • Published in last 1 year
  • In last 2 years
  • In last 3 years
  • All

Please wait a minute...
  • Select all
    |
  • Academic trend at home and abroad
    TANG Chunhua, GUO Lu, ZHANG Lili
    Journal of Diagnostics Concepts & Practice. 2025, 24(05): 485-497. https://doi.org/10.16150/j.1671-2870.2025.05.003

    In 2021, there were 93.816 million prevalent cases of stroke worldwide [age-standardized prevalence rate(ASPR) 1 099/100 000], with 11.946 million new cases in that year [age-standardized incidence rate(ASIR) 142/100 000]. Among these new cases, ischemic stroke (IS), intracerebral hemorrhage (ICH), and subarachnoid hemorrhage (SAH) accounted for 65.3% (7.804 million), 28.8% (3.444 million), and 5.8% (0.697 million), respectively. In the same year, stroke caused 7.253 million deaths, accounting for 10.7% of all global deaths. Deaths caused by IS, ICH, and SAH accounted for 49.5% (3.591 million), 45.6% (3.308 million), and 4.9% (353 000), respectively. In 2021, stroke remained the second leading cause of death worldwide, with its core disease burden indicator — disability-adjusted life years (DALYs) — exceeding 160 million, ranking third among all global total disease burdens. In terms of economic burden, the global direct medical costs and productivity losses caused by stroke reached 890 billion USD in 2021 (accounting for 0.66% of the global GDP), and are projected to exceed 1.8 trillion USD by 2050 if the current growth rate persists. The global stroke burden exhibits a dual trend of "increasing absolute numbers but decreasing age-standardized rates". Low- and middle-income countries bear most of the disease burden, and the incidence of stroke shows a coexistence of younger and older onset. In terms of risk factors, the burden of traditional behavior-related risks has decreased, while the attributable burden of metabolic and climate-related risks is rapidly increasing. China bears the heaviest stroke burden globally, characterized by a “four-high” pattern of “high incidence, high prevalence, medium-to-high mortality, and medium-to-high DALYs”, with significant urban-rural and regional disparities. This condition results from the combined effects of accelerated population aging and continuously increasing exposure to risk factors. In 2021, there were 26.335 million prevalent cases in China, with ASPR of 1 301.4/100 000. In 2021, there were 4.09 million new stroke cases in China (ASIR 204.8/100 000), accounting for 34.2% of all new global cases—far exceeding China's proportion of the world's population (about 20%). IS accounted for 67.8% [2.772 million cases, age-standardized incidence rate (ASIR) 135.8/100 000], and ICH accounted for 28.7% (1.173 million cases, ASIR 61.2/100 000). The annual total economic burden of stroke in China has exceeded 400 billion RMB, with its proportion in the national healthcare expenditure continuing to increase. Direct medical costs account for about 60%, while indirect costs (including productivity losses and caregiving expenses) account for 40%, imposing a dual pressure on both society and families. To address this challenge, a stratified precision prevention and control system centered on the coordination of "policy-healthcare-society" should be established, covering primordial, primary, and secondary prevention levels. Emphasis should be placed on cross-sector collaboration, data-driven approaches, and international experience sharing to achieve effective control of the stroke burden and promote global health equity.

  • Computing & Computer Technologies
    LIN Xiao, LU Meichen, GAO Mufeng, LI Yan
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 899-910. https://doi.org/10.1007/s12204-023-2691-y
    Human pose estimation has received much attention from the research community because of its wide range of applications. However, current research for pose estimation is usually complex and computationally intensive, especially the feature loss problems in the feature fusion process. To address the above problems, we propose a lightweight human pose estimation network based on multi-attention mechanism (LMANet). In our method, network parameters can be significantly reduced by lightweighting the bottleneck blocks with depth-wise separable convolution on the high-resolution networks. After that, we also introduce a multi-attention mechanism to improve the model prediction accuracy, and the channel attention module is added in the initial stage of the network to enhance the local cross-channel information interaction. More importantly, we inject spatial crossawareness module in the multi-scale feature fusion stage to reduce the spatial information loss during feature extraction. Extensive experiments on COCO2017 dataset and MPII dataset show that LMANet can guarantee a higher prediction accuracy with fewer network parameters and computational effort. Compared with the highresolution network HRNet, the number of parameters and the computational complexity of the network are reduced by 67% and 73%, respectively.
  • Intelligent Robots
    Li Bin, Li Zonggang, Li Haoyu, Du Yajiang
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 195-208. https://doi.org/10.1007/s12204-024-2579-5
    To optimize the movement of the three-degree-of-freedom (3-DOF) pectoral fins, a 3-DOF model of the dolphin-like pectoral fins was established, and the effects of different parameters of the pectoral fins on their propelling performance were simulated using computational fluid dynamics (CFD) technology. Using CFD simulation data as a training set and a multi-layer perceptron (MLP) neural network as a prediction model, the average thrust and lift of the pectoral fin motion under different motion cycles, rowing amplitudes, flapping amplitudes, and feathering amplitudes were predicted and modeled. A multi-objective genetic algorithm was used to obtain the optimal parameter values for maximum thrust and minimum absolute lift, and the optimal motion law for 3-DOF motion was brought. The results showed that the optimal propulsion performance was achieved at a period of 1 s, a rowing amplitude of 36 ◦ , a flapping amplitude of 18 ◦ , and a feathering amplitude of 56 ◦ . Finally, the force and displacement of the robotic fish were collected through indoor pool experiments and compared with the simulation results, indicating that the simulation results are of considerable reliability. The research results have specific guiding significance for the design of the pectoral fins of biomimetic robotic fish.
  • Computing & Computer Technologies
    YE Jihua, JIANG Lu, XIAO Shunjie, ZONG Yi, JIANG Aiwen
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 889-898. https://doi.org/10.1007/s12204-023-2688-6
    At present, research on multi-label image classification mainly focuses on exploring the correlation between labels to improve the classification accuracy of multi-label images. However, in existing methods, label correlation is calculated based on the statistical information of the data. This label correlation is global and depends on the dataset, not suitable for all samples. In the process of extracting image features, the characteristic information of small objects in the image is easily lost, resulting in a low classification accuracy of small objects. To this end, this paper proposes a multi-label image classification model based on multiscale fusion and adaptive label correlation. The main idea is: first, the feature maps of multiple scales are fused to enhance the feature information of small objects. Semantic guidance decomposes the fusion feature map into feature vectors of each category, then adaptively mines the correlation between categories in the image through the self-attention mechanism of graph attention network, and obtains feature vectors containing category-related information for the final classification. The mean average precision of the model on the two public datasets of VOC 2007 and MS COCO 2014 reached 95.6% and 83.6%, respectively, and most of the indicators are better than those of the existing latest methods.
  • Computing & Computer Technologies
    DING Leqi, WANG Biyun, YAO Lixiu, CAI Yunze
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 935-951. https://doi.org/10.1007/s12204-024-2694-3
    To overcome the obstacles of poor feature extraction and little prior information on the appearance of infrared dim small targets, we propose a multi-domain attention-guided pyramid network (MAGPNet). Specifically, we design three modules to ensure that salient features of small targets can be acquired and retained in the multi-scale feature maps. To improve the adaptability of the network for targets of different sizes, we design a kernel aggregation attention block with a receptive field attention branch and weight the feature maps under different perceptual fields with attention mechanism. Based on the research on human vision system, we further propose an adaptive local contrast measure module to enhance the local features of infrared small targets. With this parameterized component, we can implement the information aggregation of multi-scale contrast saliency maps. Finally, to fully utilize the information within spatial and channel domains in feature maps of different scales, we propose the mixed spatial-channel attention-guided fusion module to achieve high-quality fusion effects while ensuring that the small target features can be preserved at deep layers. Experiments on public datasets demonstrate that our MAGPNet can achieve a better performance over other state-of-the-art methods in terms of the intersection of union, Precision, Recall, and F-measure. In addition, we conduct detailed ablation studies to verify the effectiveness of each component in our network.
  • Automation & Computer Technologies
    LI Chunyang, ZHU Xiaoqing, RUAN Xiaogang, LIU Xinyuan, ZHANG Siyuan
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1125-1133. https://doi.org/10.1007/s12204-023-2666-z
    Bionic gait learning of quadruped robots based on reinforcement learning has become a hot research topic. The proximal policy optimization (PPO) algorithm has a low probability of learning a successful gait from scratch due to problems such as reward sparsity. To solve the problem, we propose a experience evolution proximal policy optimization (EEPPO) algorithm which integrates PPO with priori knowledge highlighting by evolutionary strategy. We use the successful trained samples as priori knowledge to guide the learning direction in order to increase the success probability of the learning algorithm. To verify the effectiveness of the proposed EEPPO algorithm, we have conducted simulation experiments of the quadruped robot gait learning task on Pybullet. Experimental results show that the central pattern generator based radial basis function (CPG-RBF) network and the policy network are simultaneously updated to achieve the quadruped robot’s bionic diagonal trot gait learning task using key information such as the robot’s speed, posture and joints information. Experimental comparison results with the traditional soft actor-critic (SAC) algorithm validate the superiority of the proposed EEPPO algorithm, which can learn a more stable diagonal trot gait in flat terrain.
  • Automation & Computer Technologies
    YU Xinyi, XU Siyu, FAN Yuehai, OU Linlin
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1085-1102. https://doi.org/10.1007/s12204-023-2631-x
    In order to solve the control problem of multiple-input multiple-output (MIMO) systems in complex and variable control environments, a model-free adaptive LSAC-PID method based on deep reinforcement learning (RL) is proposed in this paper for automatic control of mobile robots. According to the environmental feedback, the RL agent of the upper controller outputs the optimal parameters to the lower MIMO PID controllers, which can realize the real-time PID optimal control. First, a model-free adaptive MIMO PID hybrid control strategy is presented to realize real-time optimal tuning of control parameters in terms of soft-actor-critic (SAC) algorithm, which is state-of-the-art RL algorithm. Second, in order to improve the RL convergence speed and the control performance, a Lyapunov-based reward shaping method for off-policy RL algorithm is designed, and a self-adaptive LSAC-PID tuning approach with Lyapunov-based reward is then determined. Through the policy evaluation and policy improvement of the soft policy iteration, the convergence and optimality of the proposed LSAC-PID algorithm are proved mathematically. Finally, based on the proposed reward shaping method, the reward function is designed to improve the system stability for the line-following robot. The simulation and experiment results show that the proposed adaptive LSAC-PID approach has good control performance such as fast convergence speed, high generalization and high real-time performance, and achieves real-time optimal tuning of MIMO PID parameters without the system model and control loop decoupling.
  • Original articles
    WANG Yang, WANG Chao, FU Fan, ZHANG Min, LI Biao, WANG Jin
    Journal of Diagnostics Concepts & Practice. 2025, 24(05): 512-517. https://doi.org/10.16150/j.1671-2870.2025.05.006

    Objective To investigate the auxiliary value of diffuse hepatic ¹³¹I uptake (DHU) levels on post-therapy whole-body scan (Rx-WBS) images in assessing metastatic tumor burden in patients with papillary thyroid cancer (PTC) accompanied by lung metastases who underwent total thyroidectomy followed by radioiodine remnant ablation (RRA) and subsequently received ¹³¹I therapy for non-resectable distant or regional metastases. Methods A total of 22 PTC patients with lung metastases scheduled for ¹³¹I metastatic ablation therapy were retrospectively enrolled from the Department of Nuclear Medicine, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, between June 2020 and February 2025. The patients met the following three criteria: (1) total thyroidectomy; (2) completion of ¹³¹I RRA; (3) multiple pulmonary nodules detected on 131I RRA-period whole-body scan or chest CT, with stimulated thyroglobulin (sTg) >10 ng/mL. Bivariate correlation and multiple linear regression models were used to analyze the correlations of target-to-background ratios (TBR) of liver (TBRliver) and lung metastases (TBRlung) for ¹³¹I uptake with clinical parameters including sTg, thyroglobulin antibody (TgAb), and administered ¹³¹I dose. Results TBRliver showed a significant positive correlation with TBRlung (r=0.510, P<0.05). No significant correlations were found between TBRliver and sTg (r=0.218, P=0.331) or administered dose (r=0.334, P=0.128). Multiple linear regression analysis identified TBRlung as an independent influencing factor of TBRliver (β=0.511, 95% CI: 0.053-0.453, P<0.05). Conclusion In PTC patients with lung metastases after thyroidectomy and RRA, TBRliver demonstrates a significant correlation with the functional status of ¹³¹I uptake in lung metastases. Particularly when ¹³¹I scanning shows negative pulmonary nodules, elevated TBRliver may serve as an indicator of the presence of lung metastases.

  • Automation & Computer Technologies
    Wang Yan, Wang Likang, Zhang Jinfeng, Fan Xianghui
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 458-474. https://doi.org/10.1007/s12204-024-2735-y
    In recent years, underwater image enhancement techniques has received a wide range of attention from related researchers with the rise of marine resource exploitation. As the existing network feature extraction is not sufficient and the enhancement results have the problems of incomplete defogging and inaccurate color bias correction, in this paper, an underwater image enhancement method based on global dense two-branch cascade network and spatial domain grayscale transformation is proposed. The global dense two-branch cascade network can amplify the global dimensional interaction features while reducing information reduction on the one hand, and extract spatial features by obtaining spatial information at different scales to achieve richer feature extraction on the other hand; the spatial domain grayscale transformation operation can improve the contrast while color correcting the image, which makes the image visual effect better. After the training is completed, an end-to-end inference can be performed on the underwater images. The experimental results show that this paper’s model works best on the EUVP dataset, and compared with the second best, this paper’s model obtains 3.371, 0.06, 0.716, 0.024, and 1.727 improvements in PSNR, SSIM, UIQM, UCIQE, and CCF, respectively. Compared with other representative methods, the proposed network achieves significant visual enhancement in dealing with severe color bias, low light, and detail loss in underwater images.
  • Computing & Computer Technologies
    DONG Zhaoxian, YU Shuo, SHEN Yanming
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 880-888. https://doi.org/10.1007/s12204-023-2682-z
    This paper focuses on the problem of traffic flow forecasting, with the aim of forecasting future traffic conditions based on historical traffic data. This problem is typically tackled by utilizing spatio-temporal graph neural networks to model the intricate spatio-temporal correlations among traffic data. Although these methods have achieved performance improvements, they often suffer from the following limitations: These methods face challenges in modeling high-order correlations between nodes. These methods overlook the interactions between nodes at different scales. To tackle these issues, in this paper, we propose a novel model named multi-scale dynamic hypergraph convolutional network (MSDHGCN) for traffic flow forecasting. Our MSDHGCN can effectively model the dynamic higher-order relationships between nodes at multiple time scales, thereby enhancing the capability for traffic forecasting. Experiments on two real-world datasets demonstrate the effectiveness of the proposed method.
  • Automation & Computer Technologies
    TAHIR Rizwana, CAI Yunze
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1103-1113. https://doi.org/10.1007/s12204-023-2658-z
    Recent multimedia and computer vision research has focused on analyzing human behavior and activity using images. Skeleton estimation, known as pose estimation, has received a significant attention. For human pose estimation, deep learning approaches primarily emphasize on the keypoint features. Conversely, in the case of occluded or incomplete poses, the keypoint feature is insufficiently substantial, especially when there are multiple humans in a single frame. Other features, such as the body border and visibility conditions, can contribute to pose estimation in addition to the keypoint feature. Our model framework integrates multiple features, namely the human body mask features, which can serve as a constraint to keypoint location estimation, the body keypoint features, and the keypoint visibility via mask region-based convolutional neural network (Mask- RCNN). A sequential multi-feature learning setup is formed to share multi-features across the structure, whereas, in the Mask-RCNN, the only feature that could be shared through the system is the region of interest feature. By two-way up-scaling with the shared weight process to produce the mask, we have addressed the problems of improper segmentation, small intrusion, and object loss when Mask-RCNN is used, for instance, segmentation. Accuracy is indicated by the percentage of correct keypoint, and our model can identify 86.1% of the correct keypoints.
  • Automation & Computer Technologies
    SU Cheng, ZHAO Xiangtang, YAN Zengzhen, ZHAO Zhigang, MENG Jiadong
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1162-1170. https://doi.org/10.1007/s12204-023-2634-7
    Cranes used at sea have some shortcomings in terms of flexibility, efficiency, and safety. Therefore, a floating multi-robot coordinated towing system is planned to fulfill the offshore towing requirements. It is difficult to study the stability of a floating multi-robot coordinated towing system by ancient strategies. First, the minimum tension of the rope and the minimum singular value of the stiffness matrix were separately used to analyze the load stability. The advantages and disadvantages of the two methods were discussed. Then, the two stability analysis methods were normalized and weighted to obtain the method based on minimum tension and minimum singular to comprehensively analyze the stability of the load. Finally, the effect of different weighting coefficients on the load stability was analyzed, which led to a reasonable weighting coefficient to evaluate the load stability by comparing with a single analysis method. The research results provide a basis for the motion planning and coordinated control of the towing system.
  • Intelligent Robots
    Zhang Han, Zhang Guoliang, Feng Shengjie, Li Qingyun, Qu Jieming, Xie Le
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 1-11. https://doi.org/10.1007/s12204-025-2846-0
    Traditional lung biopsy procedures are complicated and time-consuming due to the lack of realtime imaging guidance, requiring physicians to frequently move between the operating room and computerized tomography (CT) imaging equipment. Robotics has been widely applied in medical surgeries, yet meeting the requirements for lung biopsy procedures with assured accuracy and safety remains a topic of research. This paper introduces a surgical robot for CT-guided lung biopsy. A kinematic analysis of the robot mechanism is conducted, and a master-slave control system tailored for this robot is developed. A force feedback algorithm is proposed to ensure the reliability and realism of the surgical process. Finally, the system’s feasibility is verified by the mechanism positioning accuracy experiment and the targeting accuracy experiment, and in vivo animal experiment is conducted to lay the foundation for clinical application.
  • Computing & Computer Technologies
    YANG Zhuang, LI Zhaofei, WANG Jihua, WEI Xudong, ZHANG Yijie
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 1065-1072. https://doi.org/10.1007/s12204-023-2675-y
    The task of identifying Chinese named entities of Chinese poetry and wine culture is a key step in the construction of a knowledge graph and a question and answer system. Aimed at the characteristics of Chinese poetry and wine culture entities with different lengths and high training cost of named entity recognition models at the present stage, this study proposes a lite BERT+bi-directional long short-term memory+ attentional mechanisms +conditional random field (ALBERT+BILSTM+Att+CRF). The method first obtains the characterlevel semantic information by ALBERT module, then extracts its high-dimensional features by BILSTM module, weights the original word vector and the learned text vector by attention layer, and finally predicts the true label in CRF module (including five types: poem title, author, time, genre, and category). Through experiments on data sets related to Chinese poetry and wine culture, the results show that the method is more effective than existing mainstream models and can efficiently extract important entity information in Chinese poetry and wine culture, which is an effective method for the identification of named entities of varying lengths of poetry.
  • Intelligent Robots
    Niu Guochen, Lü Zhihao
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 176-186. https://doi.org/10.1007/s12204-026-2900-6
    To address the technical challenge of achieving real-time and accurate detection of aerial intruders such as birds and drones in airport flight areas, where targets are extremely small, have complex and variable trajectories, suffer from strong background noise, and require long-distance detection, a tri-module fusion airspace detection network (ACE-AirDETR) based on the real-time detection Transformer (RT-DETR) framework is proposed in this paper. Performance is enhanced through three core modules. The cross-scale edge information enhancement module strengthens target contour details, generates highly discriminative features, and significantly alleviates the decline in detection accuracy caused by motion blur of small targets. The efficient additive attention module optimizes computational efficiency and improve the model’s real-time performance and deployability. The context-guided spatial feature reconstruction feature pyramid network module enhances the feature expression capability of small targets under complex backgrounds and effectively reduces the false detection and missed detection rates. To verify the effectiveness of the proposed method in specific scenarios, a self-built airplane-birddrone dataset for airspace intruders in airport-like environments is constructed. Experimental results show that compared with the RT-DETR algorithm, ACE-AirDETR improves the AP50 and AP50:95 metrics by 3.2 and 1.5 percentage points respectively, increases the frame rate by 11.8%, and reduces the computational complexity and parameter count by 20.7% and 27.3% respectively, achieving a coordinated optimization of detection accuracy, speed, and model lightweight.
  • Intelligent Robots
    Li Guolin, Chen Tong, He Shaoying, Yin Debin
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 36-47. https://doi.org/10.1007/s12204-026-2902-4
    Uncertain loads of the rigid-soft hybrid manipulator directly affect working configurations, which will alter the system model parameters, and thereby degrade control accuracy and efficiency. This paper introduces an event-triggered adaptive model predictive control strategy, which integrates with a data-driven approach to control hybrid robots with a cable-driven soft component. In the presence of model uncertainty and mismatch, adaptive identification is employed to improve the nominal model within the controller. Meanwhile, an event-triggered scheme is utilized to reduce redundant identification frequency and improve computing efficiency. Furthermore, an online data-driven method, called input mapping, uses the relationship between the historical input and output data to compensate for the minor model error in the controller via linear combination. The optimization problem is efficiently solved by designing the attenuation coefficient in an infinite-domain situation. Comparative simulation and experimental results demonstrate that the proposed method achieves improved accuracy and faster convergence speed.
  • Computing & Computer Technologies
    LIU Mengge, LIU Hao, HE Xin, JIN Shaohui, CHEN Pengyun, XU Mingliang
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 833-854. https://doi.org/10.1007/s12204-023-2686-8
    Non-line-of-sight imaging recovers hidden objects around the corner by analyzing the diffuse reflection light on the relay surface that carries hidden scene information. Due to its huge application potential in the fields of autonomous driving, defense, medical imaging, and post-disaster rescue, non-line-of-sight imaging has attracted considerable attention from researchers at home and abroad, especially in recent years. The research on non-line-of-sight imaging primarily focuses on imaging systems, forward models, and reconstruction algorithms. This paper systematically summarizes the existing non-line-of-sight imaging technology in both active and passive scenes, and analyzes the challenges and future directions of non-line-of-sight imaging technology.
  • Computing & Computer Technologies
    LIN Weiqing, LU Yanzhen, MIAO Xiren, QIU Xinghua
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 1018-1027. https://doi.org/10.1007/s12204-023-2684-x
    Self-powered neutron detectors (SPNDs) play a critical role in monitoring the safety margins and overall health of reactors, directly affecting safe operation within the reactor. In this work, a novel fault identification method based on graph convolutional networks (GCN) and Stacking ensemble learning is proposed for SPNDs. The GCN is employed to extract the spatial neighborhood information of SPNDs at different positions, and residuals are obtained by nonlinear fitting of SPND signals. In order to completely extract the time-varying features from residual sequences, the Stacking fusion model, integrated with various algorithms, is developed and enables the identification of five conditions for SPNDs: normal, drift, bias, precision degradation, and complete failure. The results demonstrate that the integration of diverse base-learners in the GCN-Stacking model exhibits advantages over a single model as well as enhances the stability and reliability in fault identification. Additionally, the GCN-Stacking model maintains higher accuracy in identifying faults at different reactor power levels.
  • Automation & Computer Technologies
    CHEN Cheng, PENG Pan, TAO Wei, ZHAO Hui
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1073-1084. https://doi.org/10.1007/s12204-023-2645-4
    Recent advances in convolution neural network (CNN) have fostered the progress in object recognition and semantic segmentation, which in turn has improved the performance of hyperspectral image (HSI) classification. Nevertheless, the difficulty of high dimensional feature extraction and the shortage of small training samples seriously hinder the future development of HSI classification. In this paper, we propose a novel algorithm for HSI classification based on three-dimensional (3D) CNN and a feature pyramid network (FPN), called 3D-FPN. The framework contains a principle component analysis, a feature extraction structure and a logistic regression. Specifically, the FPN built with 3D convolutions not only retains the advantages of 3D convolution to fully extract the spectral-spatial feature maps, but also concentrates on more detailed information and performs multi-scale feature fusion. This method avoids the excessive complexity of the model and is suitable for small sample hyperspectral classification with varying categories and spatial resolutions. In order to test the performance of our proposed 3D-FPN method, rigorous experimental analysis was performed on three public hyperspectral data sets and hyperspectral data of GF-5 satellite. Quantitative and qualitative results indicated that our proposed method attained the best performance among other current state-of-the-art end-to-end deep learning-based methods.
  • Transportation Systems
    ZHONG Ming, WU Ying, WU Chunli, WANG Fang
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1276-1288. https://doi.org/10.1007/s12204-023-2657-0
    In ports, inbound and outbound ships usually need tugboats to provide berthing and unberthing services. The decision-making problem on tugboat scheduling is important because it involves not only ships’ turnaround time at port but also tugboat operation costs. Encouraged by the problem faced by the tugboat operator, we formulate a mixed-integer programming model for tugboat scheduling problem with several practical constraints considered, such as dynamic arrival and departure of ships, qualification of tugboats, synchronization, and a flexible returning way to base to minimize the tugboat operation costs generated within the planning period. The model is inspired by genetic algorithm framework with three-dimensional coding. Effectiveness of our model and proposed solution method are testified and validated through experiments and computational results. This research helps to provide a scientific scheduling method and some insights for managers.
  • Automation & Computer Technologies
    ZHAO Xiangtang, ZHAO Zhigang, WEI Qizhe, SU Cheng
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1134-1143. https://doi.org/10.1007/s12204-023-2649-0
    Multi-robot coordinated towing system is an under-constrained system. The dynamic response of the towing system can not be fully controlled since the rope can only provide a unidirectional constraint force to the suspended object. Based on the kinematics of the multi-robot coordinated towing system with fixed-base, the Newton-Euler equations and Udwadia-Kalaba equations were used to establish the dynamics of the towing system. To obtain the motion trajectories with high stability and strong control, the motion trajectories of the towing system were optimized. During the towing, the transition from the relaxation state to the tension state of the rope was treated as a collision between the suspended object and the robot end. The trajectories of the towing system in terms of a single-variable and multiple-variable were solved, respectively. The simulation shows that the optimized trajectories are closer to reality and truly reflect the constraints of the ropes on the suspended object. The research results provide a basis for trajectory planning and control of the towing system.
  • Automation & Computer Technologies
    JIANG Yilin, ZHANG Yilong, ZHANG Fangyuan
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1114-1124. https://doi.org/10.1007/s12204-023-2654-3
    In the field of imaging, the image resolution is required to be higher. There is always a contradiction between the sensitivity and resolution of the seeker in the infrared guidance system. This work uses the rosette scanning mode for physical compression imaging in order to improve the resolution of the image as much as possible under the high-sensitivity infrared rosette point scanning mode and complete the missing information that is not scanned. It is effective to use optical lens instead of traditional optical reflection system, which can reduce the loss in optical path transmission. At the same time, deep learning neural network is used for control. An infrared single pixel imaging system that integrates sparse algorithm and recovery algorithm through the improved generative adversarial networks is trained. The experiment on the infrared aerial target dataset shows that when the input is sparse image after rose sampling, the system finally can realize the single pixel recovery imaging of the infrared image, which improves the resolution of the image while ensuring high sensitivity.
  • Automation & Computer Technologies
    HE Ximei, ZHAO Yisheng, XU Zhihong, CHEN Yong
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1220-1231. https://doi.org/10.1007/s12204-023-2624-9
    Aimed at the doubly near-far problems in a large range suffered by the remote user group and in a small range existing in both nearby and remote user groups during energy harvesting and computation offloading, a resource allocation method for unmanned aerial vehicle (UAV)-assisted and user cooperation non-linear energy harvesting mobile edge computing (MEC) system is proposed. The UAV equipped with an MEC server is introduced to provide energy and computing services for the remote user group to alleviate the doubly near-far problem in a large range suffered by the remote user group. The doubly near-far problem in a small range existing in both nearby and remote user groups is mitigated by user cooperation. The specific user cooperation strategy is that the user near the base station or the UAV is used as a relay to transfer the computing task of the user far from the base station or the UAV to the MEC server for computing. By jointly optimizing users’ offloading time, users’ transmitting power, and the hovering position of the UAV, the resource allocation problem is modeled as a nonlinear programming problem with the objective of maximizing computation efficiency. The suboptimal solution is obtained by adopting the differential evolution algorithm. Simulation results show that, compared with the resource allocation method based on genetic algorithm and the without user cooperation method, the proposed method has higher computation efficiency.
  • Intelligent Robots
    Xiao Lei, Zhao Hailong, Wu Xun, Wang Jun, Zhou Qihong
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 82-98. https://doi.org/10.1007/s12204-025-2843-3
    Industrial robots, widely employed to boost production efficiency, encounter escalating risks of joint faults as their service time lengthens. However, end-effector motion anomalies may stem from faults in the endeffector itself or from motion propagation in other joints. Moreover, the scarcity of fault samples for detection poses significant challenges. Install extra accelerometers for more precise fault diagnosis might increase the system’s complexity and costs. To tackle these challenges, this study leverages the ease of data acquisition to analyze current data from multi-joint industrial robots. A hybrid learning method is proposed for cross-device fault detection to identify the defective joint. This method integrates features from deep networks and spectral analysis to harness knowledge from both other robots and the target robot. An unsupervised model is used to assess the status of the joints based on the fused features. The proposed method’s effectiveness is validated through ablation studies and method comparisons. Results demonstrate that it accurately detects the abnormal joints without misjudgment.
  • Intelligent Robots
    Zhang Jingkai, Li Xinde, Wei Wangzichao, Wang Ziyao, Ma Ke
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 209-220. https://doi.org/10.1007/s12204-026-2903-3
    With the rapid development of unmanned aerial vehicle (UAV) technology, there is an increasingly urgent demand for intelligent detection, identification, and performance parameter inference techniques. However, existing UAV datasets face severe challenges including high acquisition costs, labor-intensive annotation, and data scarcity, and lack methods for directly predicting functional performance from single images. This paper proposes a UAV analysis framework based on synthetic data generation and multi-task deep learning. We construct an adaptive dataset generation system based on Unreal Engine 5, incorporating UAV size classification, adaptive distance adjustment algorithms, and enhanced 3D-to-2D coordinate transformation techniques for automated high-quality synthetic data generation. We design a multi-task collaborative learning network integrating visual information, distance information, and uncertainty quantification modules, supporting UAV detection, parameter prediction, functional classification, and size prediction. Experimental results demonstrate high similarity between synthetic and real data (Fr´echet inception distance is 21.05), with parameter prediction achieving mean absolute percentage errors of 29.1% and 27.0% for maximum speed and altitude respectively, and military-civilian classification accuracy reaching 62.5%. This method provides a low-cost, efficient solution for intelligent UAV analysis with considerable practical value.
  • Guideline interpretation
    DA Zhanyun, CHEN Haiye
    Journal of Diagnostics Concepts & Practice. 2025, 24(06): 613-620. https://doi.org/10.16150/j.1671-2870.2025.06.006

    Systemic lupus erythematosus (SLE) is an autoimmune disease characterized by diverse clinical manifestations, high heterogeneity, and strongly individualized treatment approaches. In 2021, the global incidence of SLE was (1.5-11.0)/100 000 person-years, while in Europe, the incidence was (1.5-7.4)/100 000 person-years. From 2009 to 2016, the incidence of SLE in the United States reached as high as 49/100 000 person-years. From 2013 to 2017 in China, analysis of the national medical insurance database and the National Rheumatology Data Center showed that the incidence of SLE in China was 14.09/100 000 person-years. Data from different countries indicate significant regional differences in the SLE incidence. With the continuous development of new diagnostic concepts and therapeutic drugs, significant progress has been made in SLE treatment strategies. However, problems such as non-standardized diagnosis and insufficient long-term management remain in the diagnosis and treatment practice of SLE in China. The "Chinese Guidelines for the Diagnosis and Treatment of Systemic Lupus Erythematosus (2025 Edition)" addresses 12 clinically relevant issues. Based on the latest domestic and international research evidence and China's SLE diagnosis and treatment practice, the guidelines provide evidence-based recommendations tailored to China's national context. These guidelines play a crucial role in promoting the advancement of standardized diagnosis and treatment and in improving the long-term prognosis of SLE patients in China. Compared with the "2020 Chinese Guidelines for the Diagnosis and Treatment of Systemic Lupus Erythematosus", the "2025 Edition" has been updated in terms of treatment targets, hormone maintenance doses, management of common organ involvement, therapeutic role of biological therapy, new immunosuppressants, and new treatment methods. This study focuses on interpreting the core recommendations of the guidelines, including SLE treatment targets, disease assessment methods, application of therapeutic drugs (including glucocorticoids, conventional immunosuppressants, and biologics), stratified treatment strategies for common organ involvement (including lupus nephritis, SLE with severe thrombocytopenia, SLE with antiphospholipid syndrome, and neuropsychiatric lupus), and long-term disease management. It aims to help clinicians quickly grasp the latest advances in SLE diagnosis and treatment, promote the implementation of standardized and individualized diagnosis and treatment concepts in clinical practice, and ultimately improve the overall diagnosis and treatment level, quality of life, and long-term survival rate of SLE patients in China.

  • Computing & Computer Technologies
    LI Shanshan, GUO Yali, HUANG Jiaxin, GAO Ruoyun
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 976-987. https://doi.org/10.1007/s12204-023-2676-x
    For traditional JPEG image encryption, block position shuffling can achieve a better encryption effect and is resistant to non-zero counting attack. However, the numbers of non-zero coefficients in the 8×8 sub-blocks are unchanged using block position shuffle. For this defect, this paper proposes a fast attack algorithm for JPEG image encryption based on inter-block shuffle and non-zero quantization discrete cosine transformation coefficient attack. The algorithm analyzes the position mapping relationship before and after encryption of image blocks by detecting the pixel values of an image by the designed plaintext image. Then the preliminary attack result of the image blocks can be obtained from the inverse mapping relationship. Finally, the final attack result of the algorithm is generated according to the numbers of non-zero coefficients in each 8 × 8 block of the preliminary attack result. Every 8×8 block position is related with its number of non-zero discrete cosine transform coefficients in the designed plaintext. It is verified that the main content of the original image could be obtained without knowledge of the encryption algorithm and keys in a relatively short time.
  • Computing & Computer Technologies
    LIU Chen, LI Wenfa, XU Yunwen, LI Dewei
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 1028-1036. https://doi.org/10.1007/s12204-023-2667-y
    The classic two-stage object detection algorithms such as faster regions with convolutional neural network features (Faster RCNN) suffer from low speed and anchor hyper-parameter sensitive problems caused by dense anchor mechanism in region proposal network (RPN). Recently, the anchor-free method CenterNet shows the effectiveness of perceiving and classifying object by its center. However, the severe coincidence false positive problem between confusing categories caused by the multiple binary classifiers makes it still insufficient in accuracy. We introduce a two-stage network CenterRCNN to take advantage of both and overcome their shortcomings. CenterRPN is proposed as the first stage to give proposals that incorporate the center keypoint idea into RPN to perceive foreground objects, replacing dense anchor-based RPN. Then the proposals are classified by the multi-classifier of RCNN header that focuses more on the difference between confusing categories and only outputs the maximum probability one of them. To sum up, CenterRPN can eliminate the drawbacks of dense anchor based RPN in Faster RCNN, and multi-classifier’s classification ability is better than that of multiple binary classifiers in CenterNet. The experiment demonstrates that CenterRCNN outperforms both basic algorithms in the accuracy, and the speed is improved as compared with Faster RCNN.
  • Intelligent Robots
    Liu Qunpo, Li Jiakun, Fei Shumin, Bo Xuhui, Naohiko Hanajima
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 106-116. https://doi.org/10.1007/s12204-025-2827-3
    For nonlinear systems with backlash-like hysteresis characteristics and external disturbance, a composite two-channel disturbance estimation adaptive controller is proposed to improve the trajectory tracking accuracy of the system. The unmodeled hysteresis and external disturbances are treated as lumped uncertainties, which are approximated by radial basis neural network and disturbance estimator respectively. These approximations are then linearly fused to form the compensation term for the lumped uncertainty. The second order linear filter is employed to estimate multiple differential terms, which are integrated into the controller design and dynamic system state updates, thereby reducing computational complexity. A weighted fusion mechanism is implemented for the two channels, and the adaptive update rate for each channel is determined based on the deviation between the lumped uncertainty reference value and the output of each channel. To address the challenges posed by the discontinuity of deviation and maintain system stability, the first-order low-pass filter is applied to smooth the deviation, enhancing system robustness. A trajectory tracking simulation of a single-input single-output nonlinear system is conducted to compare the performance of the proposed controller with baseline controllers, demonstrating the effectiveness of the composite two-channel disturbance estimation adaptive controller.
  • Computing & Computer Technologies
    ZHANG Guo, CHEN Tao, WANG Jianping
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 1037-1049. https://doi.org/10.1007/s12204-024-2723-2
    In order to meet the requirements of accurate identification of surface defects on copper strip in industrial production, a detection model of surface defects based on machine vision, CSC-YOLO, is proposed. The model uses YOLOv4-tiny as the benchmark network. First, K-means clustering is introduced into the benchmark network to obtain anchor frames that match the self-built dataset. Second, a cross-region fusion module is introduced in the backbone network to solve the difficult target recognition problem by fusing contextual semantic information. Third, the spatial pyramid pooling-efficient channel attention network (SPP-E) module is introduced in the path aggregation network (PANet) to enhance the extraction of features. Fourth, to prevent the loss of channel information, a lightweight attention mechanism is introduced to improve the performance of the network. Finally, the performance of the model is improved by adding adjustment factors to correct the loss function for the dimensional characteristics of the surface defects. CSC-YOLO was tested on the self-built dataset of surface defects in copper strip, and the experimental results showed that the mAP of the model can reach 93.58%, which is a 3.37% improvement compared with the benchmark network, and FPS, although decreasing compared with the benchmark network, reached 104. CSC-YOLO takes into account the real-time requirements of copper strip production. The comparison experiments with Faster RCNN, SSD300, YOLOv3, YOLOv4, Resnet50-YOLOv4, YOLOv5s, YOLOv7, and other algorithms show that the algorithm obtains a faster computation speed while maintaining a higher detection accuracy.
  • Computing & Computer Technologies
    LIU Biao, LIU Guangyu, FENG Wei, WANG Shuai, ZHOU Bao, ZHAO Enming
    J Shanghai Jiaotong Univ Sci. 2025, 30(5): 998-1008. https://doi.org/10.1007/s12204-023-2662-3
    Imaging sonar devices generate sonar images by receiving echoes from objects, which are often accompanied by severe speckle noise, resulting in image distortion and information loss. Common optical denoising methods do not work well in removing speckle noise from sonar images and may even reduce their visual quality. To address this issue, a sonar image denoising method based on fuzzy clustering and the undecimated dual-tree complex wavelet transform is proposed. This method provides a perfect translation invariance and an improved directional selectivity during image decomposition, leading to richer representation of noise and edges in high frequency coefficients. Fuzzy clustering can separate noise from useful information according to the amplitude characteristics of speckle noise, preserving the latter and achieving the goal of noise removal. Additionally, the low frequency coefficients are smoothed using bilateral filtering to improve the visual quality of the image. To verify the effectiveness of the algorithm, multiple groups of ablation experiments were conducted, and speckle sonar images with different variances were evaluated and compared with existing speckle removal methods in the transform domain. The experimental results show that the proposed method can effectively improve image quality, especially in cases of severe noise, where it still achieves a good denoising performance.
  • Intelligent Robots
    Zheng Luzhou, Zhao Changchen, Zhang Chao, Cheng Shichao, Zhang Jianhai
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 12-23. https://doi.org/10.1007/s12204-025-2856-y
    With the continuous advancement of sensors and algorithms, an increasing number of deep learning methods have been applied to fine-grained upper limb motion intention recognition using multimodal physiological signals. However, effectively and quantifiably integrating correlations between electroencephalogram (EEG) and electromyogram (EMG) signal channels as well as within EEG signal channels as a clue to improve performance remained challenging. In this paper, we proposed a novel framework that achieved accurate prediction of upper limb motion intentions via fusing EEG and EMG signals. Firstly, the raw input signals were fed into the feature extraction module, respectively, enabling feature decomposition in the channel dimension. Secondly, the graph convolution module with learnable edge weights was proposed to adaptively learn correlations between different modalities. Thirdly, we designed a self-attention graph pooling module that employed the self-attention mechanism to compute the attention score for each node as the basis for pooling. Compared with calculation methods using the mean or maximum value, this approach was more likely to retain nodes with stronger correlations to motor intentions. Finally, the prediction results were obtained through a classifier. We validated the effectiveness of our method on a publicly available multimodal upper limb dataset, achieving an accuracy of 93.17%.
  • Automation & Computer Technologies
    BANKOLE Adesola Temitope, IGBONOBA Ezekiel Endurance Chukwuemeke
    J Shanghai Jiaotong Univ Sci. 2025, 30(6): 1179-1187. https://doi.org/10.1007/s12204-023-2660-5
    A hybrid control strategy integrating proportional derivative (PD) and the H-infinity control methodology is proposed for a serial two-link robotic manipulator with the goal of improving the tracking performance of the robot arm. The H-infinity controller has the ability to achieve a high performance and robustness in the presence of disturbances and uncertainties, while the PD controller is effective in stabilizing the manipulator. Simulation results using Matlab and Simulink show that the proposed hybrid controller, which integrates the advantages of both PD and H-infinity controllers, has the lowest rise time for the second link, the lowest settling time for the two links, the lowest peak time for both links, and the fastest decay of the error response. In addition, the hybrid control scheme also has the lowest mean square error value, with a 53.3% improvement over the H-infinity controller and a 91.8% improvement over the PD controller, indicating an improved trajectory tracking performance when compared with pure PD and pure H-infinity controllers, respectively. It was also found that the hybrid controller has the lowest integral absolute error, integral square error, integral time absolute error, and integral time square error for the second link, while the error values for the first link are satisfactory, showing a superior performance of the hybrid controller above the PD and H-infinity controllers, respectively.
  • Intelligent Robots
    Fang Xingyu, Wei Xianming, Sun Jintao, Xu Linsen
    J Shanghai Jiaotong Univ Sci. 2026, 31(1): 99-105. https://doi.org/10.1007/s12204-026-2899-8
    In response to safety concerns caused by wafer transmission robot collisions in confined vacuum chambers, which can lead to wafer breakage and production line contamination, a collision detection and response method is proposed, leveraging the robot’s dynamics model and adjustable acceleration thresholds. The structure and motion characteristics of the robot were first analyzed. Subsequently, its kinematic and dynamic models were established and validated via simulation. To reduce noise interference, the dynamics model was computed using a reference trajectory from the host planner. Cross-correlation analysis was used to identify phase differences between the planned and encoder feedback trajectories, enabling phase compensation. For collision threshold settings, an acceleration-based adjustment scheme was developed, taking into account potential collision risk levels. Experimental tests on the vacuum robot verified the effectiveness of the proposed method.
  • Automation & Computer Technologies
    Xiao Sujie, Hao Ruipeng, Cheng Gaofeng, Xu Xiaoyan, Li Ta
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 282-288. https://doi.org/10.1007/s12204-024-2725-0
    The attention-based encoder-decoder end-to-end model has achieved promising performance in automatic speech recognition (ASR). However, in practical applications, substitution errors commonly occur in ASR systems, particularly for characters with the same or similar pronunciation. According to statistics, homophones cause at least 50% character errors. Therefore, our study focuses on addressing the issue of substitution errors with the same or similar pronunciation. In this study, we propose a BERT language model with error correction (EC-BERT) for the ASR system. We design a two-stage training schedule involving pre-training with a large amount of pseudo-paired data followed by fine-tuning with a small real-paired data to mitigate the inconsistency of the original pre-trained BERT model with our task. Unlike other error correction models, we do not need an error detection network or mask mechanism but directly use the BERT model to learn and correct the error locations. The experimental results show that our proposed method is effective and achieves a relative reduction of 19.2% in character error rate compared with the connectionist temporal classification (CTC) greedy search result and 12.8% compared with the CTC-WFST result on the AISHELL-1 test set. We also prove that our proposed EC-BERT model can achieve comparable results to other error correction models with a shorter runtime and can easily be integrated into the practical ASR system.
  • Automation & Computer Technologies
    Liu Peijin, Ding Haojian, Yan Dongyang, Sun Haofeng, Huang Tao, Li Jie
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 486-498. https://doi.org/10.1007/s12204-024-2736-x
    Aiming at the practical problems of high energy consumption and low energy efficiency during the exploitation of low-permeability oil wells because of insufficient traceability and poor matching performance of production parameters, this paper proposes a multi-objective approach for optimizing production parameters of low-permeability oil well to enhance its energy efficiency. First, a sub-model of daily liquid production yield and a sub-model of unit production energy consumption cost for single low-permeability oil well were established, and the Gaussian mixture model method was employed to compensate for the errors in the sub-model of unit production energy consumption cost, to solve the problem of the influence of uncertain facts during the oil well exploitation and to improve the precision of the model. Second, a multi-objective optimization model was established by taking into account the decision variables and constraints of the model, to maximize the daily liquid production yield while minimizing the unit production energy consumption cost. Subsequently, the non-dominated sorting genetic algorithm was employed to solve the multi-objective optimization model and obtain the production parameters. Finally, the solution set with obvious features was taken as the production parameters and applied to the actual production verification of low-permeability oil wells in a certain oil production plant of the ChangQing Oilfield. The results showed an increase in oil well production yield, and a significant energy-saving effect, thereby verifying the effectiveness of the proposed model and optimization algorithm in this paper.
  • Automation & Computer Technologies
    You Minghao, Gu Jie, Liu Shuqi
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 515-527. https://doi.org/10.1007/s12204-024-2742-z
    Deployment of integrated energy system is conducive to improving energy efficiency and achieving the transformation of the global energy system. However, recent appearance of extreme natural disasters poses a great challenge to the safe and stable operation of the integrated energy system. Therefore, the resilience of the integrated energy system, namely the ability to anticipate, withstand, respond to and recover to normal state, is to be enhanced urgently. This paper proposes a master-slave optimization model for the resilience enhancement of integrated energy system in the integrated stage of disaster response and post-disaster recovery, in view of the strong correlation between the two stages. The master model develops the optimal fault repair plan, and the sub model determines the optimal energy supply recovery scheme. Based on the master-slave model, which adopts the repair and operation state of the component as coupling variables, a coordinated optimization framework is constructed. Then, the master-slave model is merged into a two-stage robust optimization model for iterative solution in order to develop the optimal fault repair strategy and energy supply recovery scheme of the integrated energy system, enhancing its resilience in the integrated stage of disaster response and post-disaster recovery.
  • Automation & Computer Technologies
    Wang Jing, Fang Zhiqiang, Li Qianqian, Tang Zhiwei, Huang Zhangyang, Hong Zhonghua, He Haiyang
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 359-374. https://doi.org/10.1007/s12204-024-2749-5
    Urban drainage pipe system is an important part of city management. Automated detection of the status of storm drain in street-level images through current technologies in computer vision and AI is an important aspect of smart city construction. In this paper, a framework based on YOLOv5s for storm drain detection (YOLOSDD) in street view is proposed. By analyzing the characteristics of small-scale targets, YOLO-SDD focuses on optimizing the Backbone network and its loss function. Series of experiments demonstrated that in the task of detecting different states of storm drain under various environmental conditions, the mean average precision (mAP@.5) of the YOLO-SDD can reach 89.6%, increasing by 2% compared with the baseline model YOLOv5s. In the presence and absence of occlusion, the average precision of storm drain detection increased by 0.9% and 3.1%, respectively. In addition, the effectiveness and generalization ability of YOLO-SDD were further validated using the storm drain dataset of Urbana-Champaign (SDUC) from Illinois, USA, and the dataset for object detection in aerial images (DOTA). Finally, this work has deployed the YOLO-SDD on the Android system, which verifies its ability of real-time detecting storm drain in different states in street scenes.
  • Automation & Computer Technologies
    Dong Ruyi, Shi Cong
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 319-333. https://doi.org/10.1007/s12204-024-2712-5
    Accurate recognition of traffic lights is essential for ensuring the safety of passengers and pedestrians, especially in the context of self-driving car technology. However, traffic lights present challenges due to their small size and limited recognition accuracy. This paper proposes an enhanced version of the YOLOv5l algorithm specifically designed for traffic light recognition. First, the K-means++ clustering algorithm is employed to generate the prior frame. Second, the SiLU activation function in the basic convolution module is replaced with the adaptive Meta-ACONC activation function, significantly improving the model’s detection accuracy. Third, the coordinate attention mechanism is integrated into the trunk feature extraction network to incorporate coordinate information into the channel, thereby enhancing the network’s sensitivity to small target positions and mitigating the ambiguity caused by increased network depth. Finally, the network’s detection scale is improved by removing the original 20 × 20 large target detection head, leading to an improved accuracy and speed for detecting small targets. The proposed approach is evaluated on self-created traffic light datasets, and compared with the original YOLOv5l model; the improved YOLOv5l model achieves a 7.1% increase in mAP@0.5, reaching 83.3%, effectively meeting the requirements for traffic light detection and recognition.
  • Automation & Computer Technologies
    Zhao Bin, Dang Jianwu, Li Aijun
    J Shanghai Jiaotong Univ Sci. 2026, 31(2): 273-281. https://doi.org/10.1007/s12204-024-2729-9
    How neural networks coordinate to support speech perception and speech production represents a forefront research topic in both contemporary neuroscience and artificial intelligence. Despite the successful incorporation of hierarchical and predictive attributes from biological neural networks (BNNs) into artificial counterparts, substantial disparities persist, particularly in terms of real-time feedback and nonlinear regulation. To gain a more profound understanding of how BNNs manifest these attributes, the present study employed electroencephalography (EEG) techniques to examine the spatiotemporal brain network dynamics involved in listening and oral reading of identical sentences. These two tasks engage distinct sensorimotor modalities while sharing high-level semantic and syntactic representations. According to a hierarchical feedforward model, the low-level auditory and visual inputs would be progressively transformed towards abstract representations of the sentence meaning, leading to a convergence of brain network patterns in higher cognitive regions. However, our findings challenged this viewpoint by revealing an early resemblance of network activation in the prefrontal and parietal areas in both tasks. It implies a top-down predictive mechanism along with the bottom-up progression. This bidirectional interaction could be potentially implemented through frequency-specific synchronization and desynchronization between functional-specific cortical regions, laying the foundation of the speech chain system with common neural substrates.