The benchmark loads a small model to allow running on many platforms and to keep execution time low for experimentation. It will make sense to add a version to this test that loads a larger model (e.g. llama3-8b).