Species-level tree identification is a fundamental task in forest monitoring, biodiversity assessment, and climate-smart ecosystem modeling. Close-range laser scanning technologies have become indispensable tools for forest mapping because they provide high-resolution, three-dimensional structural data at the individual-tree level. However, species-level identification remains a major challenge in global environmental monitoring and AI-driven ecological assessment owing to high species diversity, structural plasticity, and variability across sensing platforms. Here, we propose the cognition-inspired multimodal attention fusion network (CI-MAFusion), a dual-branch deep learning framework that integrates point cloud data with multi-view imagery. Guided by expert dendrological reasoning and cognitive neuroscience principles, CI-MAFusion incorporates a structural branch based on an improved graph attention-based point network for encoding 3D morphological patterns and a visual branch that processes standardized multi-view projections to extract textural features. A cross-gate attention mechanism adaptively fuses structural and visual features. Each branch uses an enhanced convolutional block attention module to highlight salient features, analogous to selective attention in the human visual system. We tested CI-MAFusion using Global LiDAR TreeBank, which contains 12,057 trees from 36 species across four continents and six Köppen climate zones. The model achieved 87.50% overall accuracy at the genus level and 86.12% at the species level, outperforming unimodal and existing fusion approaches by up to 8.1%. Additionally, it further achieved > 90% overall accuracy across regions and > 80% across climate zones, with attention visualizations highlighting biologically diagnostic features such as crown contours, bark textures, and branch junctions. This cognitively inspired architecture improves generalization and advances AI-based systems toward robust recognition of biological structures in complex environments.