Research library
Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.
Search and filter
Paper list
T04. Planning, reflection, judging, and self-improving agents
Research on step-by-step reasoning, self-critique, and model-as-judge evaluation.
- T04-1Structural parallel
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- T04-4Structural parallel
Self-Refine: Iterative Refinement with Self-Feedback
- T04-7Structural parallel
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- T04-8Candidate approach
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- T04-9Similar problem
Let's Verify Step by Step
- T04-10Candidate approach
Agent-as-a-Judge: Evaluate Agents with Agents
- T04-11Similar problem
Large Language Models Cannot Self-Correct Reasoning Yet
- T04-12Candidate approach
Self-Rewarding Language Models
- T04-13Candidate approach
Voyager: An Open-Ended Embodied Agent with Large Language Models
- T04-14Candidate approach
STaR: Bootstrapping Reasoning With Reasoning
- T04-15Structural parallel
Constitutional AI: Harmlessness from AI Feedback
- T04-16Similar problem
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
From research to product use
Operating capabilities, pilots, and technologies in development are identified separately.
View technology