Stanford and Nvidia's open CLM-8B scores an agent's cached actions instead of generating tokens, cutting latency on tool ...