Aravind Srinivas

@AravSrinivas

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests
打开原帖#511482
  1. Industry

    Alexandr Wang: we are considering opening a muse merch store i am doing market resea…
  2. Industry

    Alexandr Wang: another one!
  3. Industry

    Elon Musk: Interesting