Similar Items: Truth or Mirage? 🏝 Towards End-To-End Factuality Evaluation with LLM-O asis
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Goal Alignment in LLM-Based User Simulators for Conversational AI
- Diachronic changes to the [(if the) truth BE told] construction – a corpus study
- BP-LLM : Belief Propagation for Binary Feedback in Large Language Model Alignment