Siddhant Kasture focuses on evaluating reasoning behavior in large language models, including hallucination detection, adversarial testing, and model benchmarking. At HyperQuark, he leads research on reasoning reliability and evaluation frameworks for AI systems.
Published work
Research
- Within-role, proficiency-aware discrimination on profile-derived skill graphsTwo candidates can hold the same skills at very different depths. Skill-counting cannot tell them apart. An HQ-S26 track built role graphs from real profiles and tested a score that can.
- Cross-channel disagreement as an operational diagnostic for LLM reasoning faithfulnessWhen an AI judge rates another model's reasoning, it can reward fluency over correctness. An HQ-S26 track found cases where frontier models reasoned their way to a wrong answer, and both judges rated the reasoning fully faithful.
