bytedance-research
/

UI-TARS-72B-SFT

@@ -11,18 +11,13 @@ library_name: transformers
 # UI-TARS-72B-SFT
 [UI-TARS-2B-SFT](https://huggingface.co/bytedance-research/UI-TARS-2B-SFT) &nbsp;|&nbsp;
 [UI-TARS-2B-gguf](https://huggingface.co/bytedance-research/UI-TARS-2B-gguf) &nbsp;|&nbsp;
 [UI-TARS-7B-SFT](https://huggingface.co/bytedance-research/UI-TARS-7B-SFT) &nbsp;|&nbsp;
-[UI-TARS-7B-DPO](https://huggingface.co/bytedance-research/UI-TARS-7B-DPO) &nbsp;|&nbsp;
 [UI-TARS-7B-gguf](https://huggingface.co/bytedance-research/UI-TARS-7B-gguf) &nbsp;|&nbsp;
 [UI-TARS-72B-SFT](https://huggingface.co/bytedance-research/UI-TARS-72B-SFT) &nbsp;|&nbsp;
 [UI-TARS-72B-DPO](https://huggingface.co/bytedance-research/UI-TARS-72B-DPO)
 ## Introduction
 UI-TARS is a next-generation native GUI agent model designed to interact seamlessly with graphical user interfaces (GUIs) using human-like perception, reasoning, and action capabilities. Unlike traditional modular frameworks, UI-TARS integrates all key components—perception, reasoning, grounding, and memory—within a single vision-language model (VLM), enabling end-to-end task automation without predefined workflows or manual rules.
@@ -36,6 +31,8 @@ UI-TARS is a next-generation native GUI agent model designed to interact seamles
 <!-- ![Local Image](figures/UI-TARS-vs-Previous-SOTA.png) -->
 ## Performance
 **Perception Capabilty Evaluation**
@@ -186,6 +183,7 @@ UI-TARS is a next-generation native GUI agent model designed to interact seamles
 | **UI-TARS-72B-DPO**  | **22.7** (15 steps) | - |
 | **UI-TARS-72B-DPO**  | **24.6** (50 steps) | - |
 ## Citation
 If you find our paper and model useful in your research, feel free to give us a cite.

 # UI-TARS-72B-SFT
 [UI-TARS-2B-SFT](https://huggingface.co/bytedance-research/UI-TARS-2B-SFT) &nbsp;|&nbsp;
 [UI-TARS-2B-gguf](https://huggingface.co/bytedance-research/UI-TARS-2B-gguf) &nbsp;|&nbsp;
 [UI-TARS-7B-SFT](https://huggingface.co/bytedance-research/UI-TARS-7B-SFT) &nbsp;|&nbsp;
+[**UI-TARS-7B-DPO**](https://huggingface.co/bytedance-research/UI-TARS-7B-DPO)(Recommended) &nbsp;|&nbsp;
 [UI-TARS-7B-gguf](https://huggingface.co/bytedance-research/UI-TARS-7B-gguf) &nbsp;|&nbsp;
 [UI-TARS-72B-SFT](https://huggingface.co/bytedance-research/UI-TARS-72B-SFT) &nbsp;|&nbsp;
 [UI-TARS-72B-DPO](https://huggingface.co/bytedance-research/UI-TARS-72B-DPO)
 ## Introduction
 UI-TARS is a next-generation native GUI agent model designed to interact seamlessly with graphical user interfaces (GUIs) using human-like perception, reasoning, and action capabilities. Unlike traditional modular frameworks, UI-TARS integrates all key components—perception, reasoning, grounding, and memory—within a single vision-language model (VLM), enabling end-to-end task automation without predefined workflows or manual rules.
 <!-- ![Local Image](figures/UI-TARS-vs-Previous-SOTA.png) -->
+This repository contains the model for the paper [UI-TARS: Pioneering Automated GUI Interaction with Native Agents](https://huggingface.co/papers/2501.12326).
+Code: https://github.com/bytedance/UI-TARS
 ## Performance
 **Perception Capabilty Evaluation**
 | **UI-TARS-72B-DPO**  | **22.7** (15 steps) | - |
 | **UI-TARS-72B-DPO**  | **24.6** (50 steps) | - |
 ## Citation
 If you find our paper and model useful in your research, feel free to give us a cite.