LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine Feedback.
Timon ZiegenbeinGabriella SkitalinskayaAlireza Bayat MakouHenning WachsmuthPublished in: CoRR (2024)
Keyphrases
- reinforcement learning
- function approximation
- reinforcement learning algorithms
- model free
- relevance feedback
- computer science courses
- robotic control
- flowshop
- markov decision processes
- learning process
- multi agent
- optimal control
- dynamic programming
- temporal difference
- query rewriting
- rewrite rules
- defeasible reasoning
- argumentation semantics
- learning algorithm