Improve the existing lean, constrained, and filter-aware prompt variants, and test
whether different models need different prompts to perform well.
The aim is prompts that are as general as possible, so the approach carries over to
formats that have never been seen. Format-specific prompting works but weakens the claim.
Record which models cannot do the task at all. That is a legitimate finding and belongs
in the paper.
Improve the existing lean, constrained, and filter-aware prompt variants, and test
whether different models need different prompts to perform well.
The aim is prompts that are as general as possible, so the approach carries over to
formats that have never been seen. Format-specific prompting works but weakens the claim.
Record which models cannot do the task at all. That is a legitimate finding and belongs
in the paper.