Notes from UNLP paper reviews
Cuisine
- I’m sorry, this is the longest review I’ve written, I’m honestly conflicted about this paper
- There was clearly a lot of work done, but I think it could be described much more clearly and effectively. I found myself often thinking that something I feel is a mistake actually isn’t, but it just wasn’t mentioned in the paper or it was, but I didn’t find it.
- The paper covers 3 very different points, and it’s hard to describe each of them well given the constraints of the paper.
- There are typos, missing spaces around quotes/parentheses, small details here and there — doesn’t detract from the substance of the paper, but increases the impression that the paper was either stitched together from multiple ones at the last minute, or the original wide focus didn’t leave enough time to formulate and analyze the results. I think decreasing the scope would have helped better describe the existing data.
- If the paper doesn’t get accepted, my suggestion to the authors would be to polish, focus the paper a bit more, and resubmit — the cuisine and ZNO datasets are potentially of great value to the community.
In no specific order:
- You included English/Spanish/… ZNO questions to the dataset — why? On osvita.ua I see pictures, English-language captions, and English letters (A-D instead of А-Г) — the relevance to an Ukrainian benchmark is unclear to me. The prompt used is probably in Ukrainian (which doesn’t make the non-Ukrainian questions a good fit either way), but a sample ZNO question is also missing from the Appendices, which is surprising
- Line 361, models using Russian versions of recipes during generation — if I read the prompts correctly, the only information available is the picture — how realistic is it to guess the ethnicity of the recipe based on picture alone? The “Russian Borsch” example from line 356 is problematic, because in the identification questions it explicitly conditions “you are eating in a Ukrainian restaurant”. This is absent in the generation prompt templates, just the context that the question is being asked in Ukrainian
- I don’t have an estimate how realistic it is to guess a recipe based on a picture even for a Ukrainian person in Ukraine — human evaluation data on this dataset would be crucial to contextualize the performance and errors of the models.
- It isn’t explicitly stated that the datasets will be made available and where — I’m especially interested in the full ZNO dataset w/ pictures that you collected from osvita.ua, before your filtering. I hope it is
- Table 2:
- Visual-Only VS strictly-visual % — any reason to use different words in these columns?
- “hast to” — typo
- Typos / misc:
- line 224 one ‘"’ too much
- “Appendix A” will be provided in camera-ready version — this sentence is on the page with references, making it really easy to miss if one doesn’t look for it specifically
- Table 5: different capitalization in the Qwen/quen models
- Numbers formatting — 1,234 and 1234 are both fine, but choose one
UALign
204 — intrasentential — this means basically any modifications that aren’t already listed, or it’s something more narrow?Go
EmoBench
- I’m really impressed by how much thought was put in the human annotation — passing a test, monitoring for attention, 5 independent annotations, and excluding <90% confidence answers.
- Nice table w/ licensing of all resources in the appendixes
- Solid methodology, solid paper, I don’t have many further comments honestly
Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus