Muse Spark 1.3 [LAB]
Meta's latest release of the Muse Spark model family, relevant for tracking open-weight model architectures and capabilities. https://developer.meta.com/ai/models/muse-spark/ (developer.meta.com)
BenchMIRT: What are LLM benchmarks actually measuring? [LAB]
A technical look at the validity and measurement accuracy of current LLM evaluation benchmarks. https://huggingface.co/blog/allenai/benchmirt (huggingface.co)