aboutsummaryrefslogtreecommitdiffstats
path: root/.agents/skills/ml-paper-writing/references/literature-research/paper-quality-criteria.md
blob: 0e07a8a75123a09b52ee408eaec9eeff8612b63a (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
# ML Paper Quality Evaluation Criteria

## Overview

Use these criteria to evaluate ML research papers found during literature search or when selecting papers for detailed review. The 5-dimension framework provides structured assessment for paper selection and comparison.

---

## Evaluation Dimensions

| Dimension | Weight | Description |
|-----------|--------|-------------|
| **Innovation** | 30% | Novelty and originality of contribution |
| **Method Completeness** | 25% | Clarity, rigor, and reproducibility |
| **Experimental Thoroughness** | 25% | Validation depth and analysis quality |
| **Writing Quality** | 10% | Clarity and presentation |
| **Relevance & Impact** | 10% | Domain importance and potential impact |

---

## Detailed Scoring Rubrics

### 1. Innovation (30%)

**Score 5 - Breakthrough:**
- Proposes entirely new paradigm or framework
- Solves long-standing open problem
- Major impact expected on the field
- Challenges fundamental assumptions

**Score 4 - Significant Innovation:**
- Substantial improvement over existing methods
- New insights or perspectives
- Novel combination of techniques
- Clear advancement over state-of-the-art

**Score 3 - Methodological Innovation:**
- New method or architecture proposed
- Some novelty but incremental
- Reasonable contribution
- Standard type of innovation

**Score 2 - Incremental Improvement:**
- Minor improvements to existing methods
- Limited novelty
- Small advancement
- Mostly derivative

**Score 1 - Trivial:**
- Minimal contribution
- Obvious extension
- No real innovation
- Known results

**Evaluation Questions:**
- Does this paper propose something genuinely new?
- Does it advance the state-of-the-art?
- Will this influence future work?
- Is the contribution significant or marginal?

---

### 2. Method Completeness (25%)

**Score 5 - Complete and Rigorous:**
- Full mathematical derivation
- All hyperparameters specified
- Complete algorithmic details
- Easily reproducible
- Code available

**Score 4 - Very Complete:**
- Detailed method description
- Most important details included
- Mostly reproducible
- Minor gaps in documentation

**Score 3 - Reproducible:**
- Core method clearly described
- Key details present
- Can be reproduced with effort
- Some ambiguity in details

**Score 2 - Lacks Details:**
- Key details missing
- Difficult to reproduce
- Incomplete description
- Ambiguous in important areas

**Score 1 - Unclear:**
- Method description unclear
- Missing critical information
- Cannot determine validity
- Poorly explained

**Evaluation Questions:**
- Can another researcher reproduce this work?
- Are all important details specified?
- Is mathematical derivation sound?
- Is code available and documented?

---

### 3. Experimental Thoroughness (25%)

**Score 5 - Comprehensive:**
- Multiple diverse datasets
- Extensive ablation studies
- Statistical significance testing
- Thorough analysis and discussion
- Comparison with strong baselines

**Score 4 - Very Thorough:**
- Multiple datasets
- Reasonable ablation studies
- Proper baseline comparisons
- Good analysis

**Score 3 - Adequate:**
- Main experiments complete
- Standard datasets
- Basic baselines
- Results are credible

**Score 2 - Limited:**
- Limited experiments
- Few datasets
- Weak baselines
- Minimal analysis

**Score 1 - Insufficient:**
- Minimal validation
- Toy examples only
- No meaningful comparisons
- Results not convincing

**Evaluation Questions:**
- Are experiments comprehensive?
- Are baselines strong and appropriate?
- Are statistical tests used?
- Is there ablation analysis?
- Are results on standard datasets?

---

### 4. Writing Quality (10%)

**Score 5 - Excellent:**
- Clear, precise, well-structured
- Logical flow throughout
- Professional presentation
- High-quality figures
- No ambiguity

**Score 4 - Very Good:**
- Clear and well-written
- Mostly logical structure
- Good presentation
- Minor issues

**Score 3 - Understandable:**
- Basically clear
- Some organizational issues
- Acceptable presentation
- Understandable with effort

**Score 2 - Fair:**
- Some confusing sections
- Organization problems
- Presentation issues
- Hard to follow at times

**Score 1 - Poor:**
- Unclear or confusing
- Poor organization
- Difficult to understand
- Major presentation problems

**Evaluation Questions:**
- Is the paper easy to understand?
- Is the structure logical?
- Are figures/tables clear?
- Is the writing professional?

---

### 5. Relevance & Impact (10%)

**Score 5 - High Impact:**
- Solves important problem
- Broad applicability
- Expected wide influence
- Addresses fundamental challenge

**Score 4 - Domain Important:**
- Important problem in field
- Significant potential impact
- Relevant to many researchers

**Score 3 - Meaningful:**
- Meaningful contribution
- Moderate impact expected
- Relevant to subset of field

**Score 2 - Niche:**
- Specialized problem
- Limited applicability
- Narrow impact

**Score 1 - Limited:**
- Very narrow problem
- Minimal impact expected
- Limited relevance

**Evaluation Questions:**
- Is this an important problem?
- Will this influence future work?
- Is it relevant to current research needs?
- Does it address a significant challenge?

---

## Scoring Calculation

**Weighted Total:**
```
Total = (Innovation × 0.30) + (Method × 0.25) + (Experiments × 0.25) + (Writing × 0.10) + (Impact × 0.10)
```

**Example Calculation:**
- Innovation: 4/5
- Method: 3/5
- Experiments: 4/5
- Writing: 3/5
- Impact: 4/5

```
Total = (4 × 0.30) + (3 × 0.25) + (4 × 0.25) + (3 × 0.10) + (4 × 0.10)
      = 1.20 + 0.75 + 1.00 + 0.30 + 0.40
      = 3.65 / 5.0
```

---

## Selection Process

### For Literature Reviews

1. **Screen papers** by title/abstract for relevance
2. **Full review** of potentially relevant papers
3. **Score each paper** using all 5 dimensions
4. **Rank by total score**
5. **Select top papers** for detailed review

### Quality Thresholds

- **Excellent**: 4.0+ (include definitely)
- **Good**: 3.5-3.9 (include if relevant)
- **Fair**: 3.0-3.4 (include if highly relevant)
- **Poor**: <3.0 (exclude unless essential)

---

## Quick Screening Indicators

Before detailed review, check:

**Positive Indicators:**
- Published at top venue (NeurIPS, ICML, ICLR)
- Citations in top papers
- Code available with stars
- Authors from top labs
- Clear novelty in abstract

**Negative Indicators:**
- Vague abstract
- Limited experiments mentioned
- No baselines mentioned
- Poor writing in abstract
- incremental claims only

---

## Integration with Paper Discovery

When using arXiv search (`arxiv-search-guide.md`):

1. **Search** for relevant papers
2. **Extract metadata** from arXiv pages
3. **Quick screen** by abstract/relevance
4. **Detailed review** of promising papers
5. **Score using** these criteria
6. **Rank and select** top candidates

---

## Notes

- These criteria are designed for ML papers specifically
- Adjust weights based on your specific needs
- Use scores as relative comparisons, not absolute judgments
- Consider venue reputation as additional signal
- Code availability is increasingly important for reproducibility