RESEARCH Poolside lanca Laguna S 2.1, modelo open-weight 118B superando DeepSeek V4 Flash em tarefas agentísticas 25 de jul., 04:34 · globenewswire.com
RESEARCH Kimi K3 atinge 68.5% em DeepSWE, custa 2.8x menos por tarefa que Claude Fable 5 24 de jul., 20:34 · together.ai
RESEARCH LangChain publica framework de benchmark Deep Agents; três suite eval (Harbor-Index, τ³-bench, ContextBench) estabelecem padrão de autonomia de agente longo-horizonte ontem · langchain.com
RESEARCH Black Forest Labs FLUX 3: modelo multimodal unificado gera vídeo, áudio e ações robô de arquitetura única ontem · manilatimes.net