Voxel-MAE for subsurface geology
Single-stride sparse transformer with a voxel-MAE objective used directly for dense regression on near-field mineral exploration data.
Near-field mineral exploration is the problem of predicting where ore sits in the volume immediately around an existing drilling network: the 100 m to a few km of subsurface that is geologically continuous with a working mine but only sparsely sampled. The data is unforgiving - drill holes are tens to hundreds of metres apart, the measurements they yield (assays, mineralogy, geotechnical logs, sulphidation, distance to fault planes) are multimodal and rarely co-located, and the variables of interest, copper and gold grade, are heavily right-skewed, with economically viable ore making up only a tiny fraction of any volume.
This work, a collaboration with Maptek on data from Dundee Precious Metals’ Chelopech copper-gold mine in Bulgaria, treats the borehole network as a 3D point cloud and trains a model to fill in grade everywhere else. Maptek’s GeologyCore pipeline produced the upstream feature set: drill holes segmented into 1 m intervals, each interval flagged by triangulation against the existing fault-and-orezone wireframes, tagged with the fault block it sits in, and annotated with distance to the nearest fault plane. After preprocessing — filtering to ~920k sample points with non-missing Cu/Au, dropping the lithology one-hots (which empirically hurt convergence), and applying reversible log (Cu) and Box-Cox (Au) redistributions to flatten the grade skew — the input is a 23-feature, georeferenced point cloud over the active mining region.
The model is a Single-Stride Sparse Transformer backbone with a Voxel-MAE masking objective. The volume is voxelised at 5 × 5 × 3 m and partitioned into 150 × 150 × 60 m blocks; within each block, an 8 × 8 × 8 sliding window defines the receptive field for sparse regional attention with region-shifting between layers, so attention is computed only over non-empty voxels. Voxel-MAE is normally used as a pretraining objective — mask some non-empty voxels, predict the inputs, then fine-tune a downstream head — but here it is the downstream task: an encoder + decoder + regression head trained end-to-end to mask 40% of non-empty voxels at random and predict redistributed Cu or Au grade at the masked positions under MSE loss. Ablations on masking ratio, voxel size, network depth, and window aspect ratio selected the final 40% / 5 m × 5 m × 3 m / 5-encoder + 5-decoder / 8 × 8 × 8 configuration.
On held-out blocks of the Chelopech mine, the final models reach NRMSE 0.05 on copper and 0.03 on gold, with class-weighted F1 of 0.68 and 0.54 across four economic-viability bands. On the minority high-opportunity class — the class that actually matters for exploration decisions — F1 is 0.21 (Cu) and 0.12 (Au), and precision on the high-gold class is 0.59: when the model predicts a high-grade gold voxel it is usually correct, even though recall is bounded by the class imbalance.
Inference is iterative: at each sample block the model is asked to predict at randomly seeded unobserved voxels, the most confident predictions are appended to the context, and the process repeats until the block reaches a target density. Aggregating across all blocks and ranking by combined Cu/Au signal yielded five prioritised exploration targets, the strongest of which sat at 7383 ppm Cu and 1.22 ppm Au in a part of the mine well away from existing extraction. The work was recognised with a $4,000 Future Explorers innovation award and feeds into 3D subsurface modelling work at New Gradient.