The Educational Designer

judgement scale and gavel in judge office

Navigating AI generated assessment and academic integrity standards

At the beginning of 2023, the landscape of higher education changed when ChatGPT became known to the community. Since then, many generative AI presentations have revolved around academic integrity standards. It’s shaking up the industry because it’s making us rethink what constitutes ‘student work’ and ‘cheating’. A tool could be used to formulate ideas, troubleshoot scenarios and evaluate and give feedback on work, but then… so can humans. It’s not a new risk or a new concept at all. A student can submit an essay written entirely by generative AI, or they could submit an essay written entirely by an essay writing service. The ‘threat’ is not the tool itself, and that threat will never go away; it will just evolve.

Can educators create assessments that can defend against AI generated text? 

The simplest answer to this question is that if a student can write an assessment submission using generative AI with very little effort involved and pass, then the issue may not be generative AI. If they are driven to do this, then we may need to work on the student ‘buy-in’ when it comes to developing these skills for their future professions. Consider the parameters around the assessment:

Does higher-order thinking move students away from AI-generated submissions?

There are many situations in which we need to test lower-order thinking. Basic factual recall is an important task in many professions, and we expect students to have certain knowledge ingrained before they can perform more complex tasks. This kind of knowledge testing is valid and can be invigilated, but this ingrained knowledge can also show up in other tasks. Any assessment regime needs to have a decent percentage of higher-order thinking.

Create scenarios with multiple layers of complexity that ask students to analyse and evaluate. 

Present students with scenarios that require analysis, synthesis, evaluation, and personal insight. For example, an assessment task could incorporate the perspectives of local stakeholders. Students should consider how various stakeholders, such as government bodies, community organisations, or industry players, might respond to a given situation based on local regulations. To include diversity in assessment, have students select different local contexts, and have a group component of the assessment that requires them to compare and contrast their findings. Embed culturally and contextually relevant details to make them more representative of their actual context. 

Require students to integrate local policies, procedures, and frameworks in their responses and connect the scenario to the literature.

Have students draft mock proposals or action plans that outline how they would practically implement a national health framework in a specific local context. Have the students support their rationale for the local context with evidence from the literature. 

Structure assessments to include multiple stages, so that the results of one stage must be carried through and integrated into the next stage.

Give the same scenario for different contexts, and have students highlight the different levels of success a project or initiative could have based on the constraints of that context. Then, identify in the literature evidence that could show a Western democratic bias. Finally, have students seek literature that is specific to the cultural context and compare the outcomes. 

Are we overly penalising AI generated text?

This comes with the caveat that generative AI consumes a significant amount of energy, and the majority of the LLM base is created with text that values white, male, heterosexual cisgender voices. Generative AI is a tool, but one that must be used with an understanding of the ethical implications.  Any tool can be used and misused depending on the student’s knowledge and practice. A more fruitful outcome would be to teach students how to use the tool in a way that supports their own critical thinking.

How can we tell when a student has just submitted AI generated text verbatim?

Some institutions use AI detection software; others don’t. The key here is that AI detection software is not completely accurate, has been known to generate false positives, and has been identified as being biased against non-native English speakers (Liang et al., 2023). Consider the impact of a false accusation on a student who has submitted their own work only to be accused of cheating. Particularly a student who is in a country that is different from their own and uses a language that is not native to them. These tools should be used as a data point that can guide analysis and conversations, and we need to ensure we protect our marginalised students. 

Generative AI does have a somewhat visible footprint, in the form of word choice, sentence and idea structure. However, this is getting more nuanced over time as humanity reacts to words like ‘surfacing’, and ‘delve’, and devices like the em dash. Compare the article you have just read to the conclusion below. Generative AI was used in the formation of this article; in particular, in the structure and ideas. Basically, as an ideation tool. However, a significant amount of contextually appropriate experience and information was needed to create the article you have read. Compare this to the conclusion that ChatGPT generated for me:

In Conclusion: Navigating the AI Frontier with Integrity

As technology blurs the lines between human and artificial intelligence, the responsibility of educators to safeguard academic integrity becomes even more critical. The key is not to resist the technological tide but to ride it with a firm grip on ethical values. By setting clear guidelines, crafting thought-provoking questions, and nurturing students’ unique perspectives, we can ensure that assessments remain a true reflection of their intellectual journey. So, let’s embark on this adventure, armed with knowledge, innovation, and an unwavering commitment to maintaining the sanctity of education in the age of ChatGPT.

It follows a common pattern. “As blah happens, blah becomes more important. The key is not blah, but blah. Let’s delve into blah. By doing blah, blah, and blah, we can blah. So let’s blah weasel words blah. Balh blah call to action.”

If it sounds like an edtech startup trying to sell its product (and you die inside a little when you read it), it may be written by generative AI.

References

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779

Get new blog posts to your inbox.

Discover more from The Educational Designer

Subscribe now to keep reading and get access to the full archive.

Continue reading