When starting training, shut down the inference subprocess first so the training subprocess has full GPU memory available.