Presentation Information

[O12-P55]A Multidimensional Behavior Simulation for Autonomous Acquisition of Dinosaur Gait Cycles Using Reinforcement Learning

*Hiroyuki Tanaka1 (1. Shizuoka Prefectural Hamamatsu Kita High School)

Keywords:

Dinosaur,Functional morphology,Computer simulation

1. Background Behavioral analyses of dinosaurs based on body fossils are generally considered less reliable than trace-fossil approaches, and simulation-based gait reconstructions have rarely been attempted. A prior study by Naruse et al. at Kitami Institute of Technology used a Central Pattern Generator (CPG) to produce a bipedal gait cycle, adjusted by an Artificial Neural Network (ANN) and Genetic Algorithm (GA). However, the simplified 3D model lacked digitigrade posture, a key anatomical feature of dinosaurs, and the system did not allow direct joint-level control, limiting biological plausibility.

2. Objective This study proposes an End-to-End (E2E) locomotion control simulation using reinforcement learning without hand-crafted features. An agent independently controls each joint of a dinosaur 3D model whose biophysical parameters are derived from body fossils, and learns control patterns that achieve stable posture and gait through reward maximization. The resulting gait is discussed in terms of functional morphological validity.

3. Data A 3D model of Parasaurolophus, a Late Cretaceous hadrosaur from North America, was used for bipedal locomotion acquisition. Based on prior studies, the total length was set to 9.45 m and mass to 3 t. Fifteen independently controllable joints were configured: 3 in the neck, 1 in the torso, 4 per hindlimb (8 total), and 3 in the tail. The simulation environment employed Unity, PyTorch, and ML-Agents.

4. Methods Proximal Policy Optimization (PPO) was used as the learning algorithm. The agent was spawned on a 50 m square field and required to reach a target placed 7-15 m ahead via leg control. Contact between the ground and any part other than the plantar surface or tail tip was classified as a fall and penalised. A bootstrapped curriculum progressively increased reward and penalty magnitudes across four phases from lenient to full strength to enhance learning efficiency and policy stability. Controllable joints were restricted to the legs and vertical tail motion; neck joints are rigidly locked.

5. Results The agent successfully acquired a stable walking gait that consistently reached the target. The locomotion exhibited a slight tendency to drag its feet but maintained low head oscillation, appropriate stride length, and steady velocity. Cumulative reward dipped at each phase transition but showed sustained growth. The mean episode duration for successful episodes was 4.8 s, corresponding to approximately 1.58 km/h.

6. Discussion The agent autonomously acquired smooth gait and stable posture. Stride length and velocity emphasized in the bootstrap curriculum were clearly reflected in the corresponding phases, and the achieved velocity closely matched the 1.5 km/h target. These results suggest that the system is effective for acquiring stable dinosaur locomotion and that the resulting gait likely possesses biological plausibility. Future work includes turning and sequential target reaching, validation with extant digitigrade animal models, and footprint formation simulation.

7. Acknowledgments This research was supported by the Future Scientist School (FSS) program at Shizuoka University. The author gratefully acknowledges Prof. Toru Aoki and Asst. Prof. Yuki Kase of the Aoki-Kase Laboratory, Faculty of Informatics, Shizuoka University.

8. References Naruse et al., "Acquisition of Walking Behavior of Dinosaurs in a 3D Physical Space," Proc. JSME Conf. Robotics and Mechatronics, 1P1-M05 (2013). Lull, R.S. & Wright, N.E., Hadrosaurian Dinosaurs of North America. GSA Special Papers 40 (1942). Paul, G.S., The Princeton Field Guide to Dinosaurs, 2nd Ed. Princeton Univ. Press (2016). Juliani, A. et al., "Unity: A General Platform for Intelligent Agents," arXiv:1809.02627v2 (2020).