Research library

Search papers on equipment connectivity, quality inspection, production analysis, and manufacturing work.

Search and filter

Paper list

T04. Planning, reflection, judging, and self-improving agents

Research on step-by-step reasoning, self-critique, and model-as-judge evaluation.

  1. T04-1Structural parallel

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

    Authors
    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou
    Year
    2022
    Venue
    NeurIPS 2022

    Korean review of T04-1

  2. T04-4Structural parallel

    Self-Refine: Iterative Refinement with Self-Feedback

    Authors
    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark
    Year
    2023
    Venue
    NeurIPS 2023

    Korean review of T04-4

  3. T04-7Structural parallel

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

    Authors
    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, Ion Stoica
    Year
    2023
    Venue
    NeurIPS 2023 Datasets and Benchmarks Track

    Korean review of T04-7

  4. T04-8Candidate approach

    G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

    Authors
    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, Chenguang Zhu
    Year
    2023
    Venue
    EMNLP 2023

    Korean review of T04-8

  5. T04-9Similar problem

    Let's Verify Step by Step

    Authors
    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T04-9

  6. T04-10Candidate approach

    Agent-as-a-Judge: Evaluate Agents with Agents

    Authors
    Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, Jürgen Schmidhuber
    Year
    2024
    Venue
    arXiv

    Korean review of T04-10

  7. T04-11Similar problem

    Large Language Models Cannot Self-Correct Reasoning Yet

    Authors
    Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T04-11

  8. T04-12Candidate approach

    Self-Rewarding Language Models

    Authors
    Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, Jason Weston
    Year
    2024
    Venue
    ICML 2024

    Korean review of T04-12

  9. T04-13Candidate approach

    Voyager: An Open-Ended Embodied Agent with Large Language Models

    Authors
    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar
    Year
    2024

    Korean review of T04-13

  10. T04-14Candidate approach

    STaR: Bootstrapping Reasoning With Reasoning

    Authors
    Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman
    Year
    2022
    Venue
    NeurIPS 2022

    Korean review of T04-14

  11. T04-15Structural parallel

    Constitutional AI: Harmlessness from AI Feedback

    Authors
    Yuntao Bai, Saurav Kadavath, Sandipan Kundu and others, 51 authors in total (Anthropic)
    Year
    2022
    Venue
    arXiv

    Korean review of T04-15

  12. T04-16Similar problem

    CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

    Authors
    Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, Weizhu Chen
    Year
    2024
    Venue
    ICLR 2024

    Korean review of T04-16

From research to product use

Operating capabilities, pilots, and technologies in development are identified separately.

View technology