검색 상세

Effects of Raters, Task Types, and Rating Rubrics on Writing Assessment in EFL Context

초록/요약

The primary purpose of this study was to investigate the effects of raters, task types, and rating rubrics on the assessment of writing ability in the EFL context. These variables have been recognized as the three most important variables that influence language writing assessment. Thus, the interaction between each of these variables and the evaluation of the writing performance of the subjects in this study was carefully examined. Specifically, the effects of the two rater groups and the comparability of the tasks and rating rubrics on the assessment of writing ability were researched. A total of 62 narrative and persuasive essays, written by intermediate level college students, were rated by two groups of professors, one native-speaking (NES) and the other non-native speaking (NNS), on the basis of both holistic and analytic rubrics. The holistic rubric was comprised of a 5-point scale where each level of points matched a description that reflected the level of achievement in the following areas: content control, organization control, socio-pragmatics control, grammar control, and mechanics. The analytic rubric utilized 5 subscales of the same domains as the holistic rubric, and the examinees’ writing performance was rated on all 5 subscales. The data were analyzed using analytic methods such as multivariate generalizability theory and multi-faceted Rasch model. The level of rater severity was examined and sources of variance in each of the tasks were identified. Also, the tasks were examined to see whether they were actually measuring the language domains that were described by the rating rubrics. The results showed that the test takers’ scores on the writing test did not exhibit significant differences when rated by two different groups, namely, the NES and NNS raters. However, the results indicated that the test performance was affected by the writing tasks, the rubric, and the subscale domains. The tasks were found to be non comparable when rated by means of both rubrics. More variance was found in the narrative writing task than in the persuasive writing task. Even using the same language domains described in the rubric, the evaluation of the narrative writing carried with it the raters’ own judgment of the writing task, a judgment that seemed to go beyond the guidelines described in the rating rubric. In addition, where content was found to be the most important predictor variable in the narrative writing task, grammar was found to be most important predictor variable in the persuasive writing task. Socio-pragmatics was found to have the most bias interaction with the raters, the least correlation with other variables and was the least important predictor variable in both tasks. Suggestions for improving the quality of writing assessment include the following: more focused training for raters, especially with a view to awareness raising of the raters’ own severity level; using analytic rubrics when placing students into proficiency levels; designing and developing different rubrics for different task types and assessment purposes; and increasing reliability by employing a greater number of writing tasks.

more

목차

CHAPTER1 INTRODUCTION
1.1 Background of the Study 1
1.2 Research Questions 4
CHAPTER2 REVIEW OF THE LITERATURE
2.1 L2 Writing ability 6
2.2 Factors Contributing to Variability in ESL/EFL Essay Rating 10
2.2.1 The Rater 10
2.2.2 The Writing Task 13
2.2.3 The Rating Method 17
2.3 Review of Analytic Methods 23
2.3.1 Generalizability Theory 23
2.3.2 The Rasch Model 25
CHAPTER3 METHODOLOGY
3.1 Participants 28
3.1.1 Examinees 28
3.1.2 Raters 29
3.2 Instruments 30
3.2.1 Writing Prompts 30
3.2.2 Rubrics for Assessing L2 Writing 31
3.3 Data Collection and Procedures 32
3.3.1 Writing Sample Collection 32
3.3.2 Rating Procedures 32
3.3.3 Statistical Procedures 33
CHAPTER4 RESULTS AND DISCUSSION
4.1 Descriptive Statistics 36
4.2 FACETS Analysis Using Holistic Rating Method 42
4.2.1 Facets Summary 42
4.2.2 Examinee Ability Measure 45
4.2.2.1 Measurement Report of Examinee Abilities 46
4.2.2.2 Variability of Examinee Ability Measures 49
4.2.2.3 Reliability of Examinee Ability Measures 50
4.2.2.4 Unexpected Responses for Examinees, Tasks, and Raters 51
4.2.3 Rater Severity 52
4.2.4 Task Types 56
4.2.5 Rating Rubric 57
4.2.6 Bias Analysis 61
4.2.6.1 Bias Analysis of Raters across Examinees 61
4.2.6.2 Bias Analysis of Raters across Task Types 67
4.3 FACETS Analysis Using Analytic Rating Method 70
4.3.1 Facets Summary 70
4.3.2 Examinee Ability Measure 72
4.3.2.1 Measurement Report of Examinee Abilities 74
4.3.2.2 Variability of Examinee Ability Measures 75
4.3.2.3 Reliability of Examinee Ability Measures 75
4.3.2.4 Unexpected Responses for Examinees, Tasks, and Raters 76
4.3.3 Rater Severity 78
4.3.4 Task Types 79
4.3.5 Rating Rubric 80
4.3.6 Bias Analysis 85
4.3.6.1 Bias Analysis of Raters across Examinees 85
4.3.6.2 Bias Analysis of Raters across Task Types 90
4.3.6.3 Bias Analysis of Raters across Rating Rubrics 93
4.4 Multivariate Generalizability Analysis 95
4.4.1 Generalizability of Scores for the Holistic Rating Session 96
4.4.1.1 Analysis of Variance 96
4.1.1.2 G-Study 97
4.4.1.3 D- Study 99
4.4.1.4 G-Facets Analysis 102
4.4.2 Generalizability of Scores for the Analytic Rating Session 103
4.4.2.1 Analysis of Variance 103
4.4.2.2 G-Study 104
4.4.2.3 D-Study 107
4.4.2.4 G-Facets Analysis 110
4.5 Regression Analysis 111
4.5.1 Analysis for Narrative Writing Task 111
4.5.2 Analysis for Persuasive Writing Task 114
CHAPTER5 CONCLUSION AND IMPLICATIONS
5.1 Summary of the Results 118
5.2 Implications of the Study 125
5.3 Limitation of the Study 128
REFERENCES 131
APPENDICES 138

more