What is effort really? When do you change it it and why not just use max effort for everything? I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results. https://t.co/KO2D51j34H
Posts de pesquisadores, desenvolvedores e lideranças de IA no X.
29 posts disponíveis
What is effort really? When do you change it it and why not just use max effort for everything? I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results. https://t.co/KO2D51j34H
Manually turned off all the AI features. Survived approximately five minutes without autocomplete. Realizing now that what I want isn't "removal of AI", it's removal of UI clutter and prompt boxes. https://t.co/GN15SliDMZ
We are back to daily Claude updates. And I’m all in for it. Nice update! https://t.co/4udoimbKYy
There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to. We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations. We are prioritizing as best as we can based on severity, and adding resources. Hugging Face is still the most severe event we’ve seen. We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.
Opus 5.5 got my ts-rust port working in 10 hours, and it has been grinding on performance for the last 24 hours. I've been working on this port on and off for about 4 months. I got to ~35% tests passing with GPT-5.6 Sol, and ~85% with GPT-6 Astra. Both models stalled hard once hitting those numbers, and ran in loops with no meaningful progress for days at a time. I've never had "enough Anthropic tokens" to try a Claude model on a port like this. Opus 5.5 feels practically unlimited, so I threw it a "/goal finish the port and make it faster". I can not believe how quickly Opus unblocked the work Astra was stuck on. It may have made this port an actually viable project. Absolutely mind blown right now.
What are specific requests in Claude Computer/Browser Use that have failed for you? We want to fix. An example for me is: "pay my barber $40 through my personal paypal" (where paypal is logged into my 2nd chrome profile, I'm logged out but password is saved)
We've now received GPT-6 Astra and Opus 5.5. Confirmed for the next releases are: - Sonnet 5.5 - Haiku 5.5 - Kimi K3.1 (they themselves hinted at it in a cryptic post) - Gemini 4, before the end of the year Unconfirmed but very likely: - Fable 5.5 - Astra update (at the DevDay?) Currently, there's no word yet on new GLM or DeepSeek models.
New on the Science Blog: Yes, Claude can do Nine Loops. Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators. Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access? Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog. Read more: https://t.co/CS2f2qoIhJ
Tag escreve >50% dos meus PRs todos os dias. Ele também faz ~100% da minha análise de dados e corrige a maior parte do feedback do produto + bugs. Não é como um bot comum do Slack. Ele é proativo, programável, tem memória, tem acesso aos seus conectores. Com Opus 5.5 e Fable 5.1, ele tem um forte julgamento. Exemplos de prompts que uso: - "@Claude reaja a threads neste canal com ✅ quando resolvido" - "@Claude a partir de agora, tente reproduzir todos os bugs neste canal do início ao fim, executando o aplicativo completo. assim que reproduzir o bug, envie um pr para corrigi-lo e marque a equipe certa para revisão de código" - "@Claude crie ~100 hipóteses sobre o que pode explicar esses dados estranhos, use um fluxo de trabalho para validar/invalidar. gaste cerca de 10 milhões de tokens explorando muito a fundo. desenhe um gráfico com o resultado." - "@Claude crie um jogo interativo explicando como esta parte do código funciona, depois crie um slide deck explicando para outras pessoas da equipe" https://t.co/gmsALM8Lvs
GPT-6 Astra ajuda a Proaction a construir agentes de gerenciamento de frota mais rapidamente, com execuções de uso de computador mais curtas para realizar o mesmo trabalho. https://t.co/74fsl9vgvM
CHATGPT, POR FAVOR NÃO ME COMA VIVO https://t.co/mCHEquroBr
Trump diz que Xi gostou da ideia de chamar a IA de “super inteligência”. Porque isso seria “SUPER!” Todos concordaram. Eu não usei superinteligência para escrever esta postagem. https://t.co/5IcUdiGZqD
Eu adoro como Mark Zuckerberg simplesmente ri da pergunta se "a IA vai nos exterminar a todos", como se fosse a pergunta mais absurda do mundo. A essa altura, ninguém vai desacelerar. https://t.co/Iglv5IW0qP
Então o plano "ChatGPT Pro max" por US$ 500 não só incluirá taxas mais altas, mas também inferência mais rápida. „Trabalho mais rápido em codex“ https://t.co/JuMI9ibFOQ
Para aqueles que estão voltando para o Claude Code depois de usar o aplicativo ChatGPT/Codex por um tempo, como vocês se sentem sobre a mudança? O que vocês sentem falta?
Tenho certeza de que o fornecimento de energia, as conexões de rede e a construção de data centers estarão entre os maiores gargalos da IA nos próximos anos. O Projeto Suncatcher do Google explora uma possibilidade interessante: mover a computação de IA para a órbita, com satélites equipados com TPUs conectados por lasers. O Google diz que painéis solares na órbita certa poderiam ser até 8 vezes mais produtivos do que na Terra, graças à luz solar quase contínua. Parece loucura, mas o potencial é enorme: escalar a computação sem esperar por conexões de rede terrestres seria uma grande vantagem. Resfriamento, custos de lançamento e conexões confiáveis entre satélites ainda precisam ser resolvidos. O primeiro passo é um protótipo desenvolvido com a Planet, programado para ser lançado a bordo da missão Transporter-18 da SpaceX. A SpaceX pode acabar sendo uma das maiores vencedoras aqui. Tornar a computação orbital econômica exigiria lançamentos acessíveis e frequentes. Isso poderia transformar o boom da infraestrutura de IA em uma nova fonte importante de demanda para a SpaceX.
A maioria das pessoas pensa no T3 Code como uma alternativa ao Codex. Eu penso nele como uma alternativa ao Xcode. https://t.co/Caua83ToRv
É aqui que estávamos com o código de IA no início do ano passado, a propósito https://t.co/UwHGZYeJX5
Acabei de falar com um amigo meu que era super contra IA mas depois de experimentar o Opus 5.5 ele agora é 'AI pilled' e gosta de 'vibe coding'
1/ Octen gives AI agents live web search at 0.21 seconds per query in Artificial Analysis' benchmark. It is the only Search API in the top 3 for quality, speed and search cost. Search fees: roughly 1/7 of Exa auto's. Sept 8 data. @OctenAI @KZouAPT https://t.co/1kTISdMQ6l
Damn, im in love with Opus 5.5. Its the best model ive ever used, has such a good taste, is comparatively quickly, smart and finally precise instead of verbose. As good as Opus 5.5 is, I'm incredibly excited for Fable 5.5. After switching completely to Codex, I now have to go back to Anthropic. oh and btw. its crazy how fast the vibe has changed. Now all I read on X are posts about how great Opus tastes and how OpenAI absolutely has to catch up. One day that changed everything.
The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge https://t.co/nXvJanC5NT
Anybody have a VS Code fork with all of the extra features and AI stuff removed? I want a minimal editor that doesn't spam me with requests to use a mediocre agent when I just want to quickly look at some code changes.
Anthropic reportedly plans to give its seven co-founders majority voting control ahead of its IPO, which gives them much more power. The Information reports that a proposed Palantir-style share structure would give them 50.1% of voting power on most corporate matters, as long as at least three founders retain a minimum number of shares. Dario Amodei reportedly owns around 2% of Anthropic. The new shares would add voting power without additional economic ownership.
incredibly thoughtful piece on AI and creativity hopefully we can make our producst better at enabling and amplifying creatives https://t.co/SeI25wbaI7
It's insane how much better the $200 Claude Code plan is compared to Codex right now. Just a few weeks ago, it was the other way around. Excited to see how OpenAI fights back.
Opus 5.5 just filed a pull request that changes the streaming behavior in T3 code. It's so cool that I can send off a prompt with a vague idea of what I want, and the result is a ready-to-merge pull request with a video demo of the changes directly in the PR. The PRs I make with AI are significantly better than the ones I used to make by hand. The future is awesome.
Fiz todos os modelos novos e sofisticados fazerem uma grande auditoria para mim, depois montei um painel de juízes para classificar a qualidade de suas auditorias. Astra foi massacrada. Opus e Fable logo atrás. Grok 4.7 e GPT-6 Sol ficaram MUITO para trás. https://t.co/r4g6zHSG8w
Novo plano profissional da OpenAI chegando. Custará US$ 500. No entanto, ainda não há informações sobre as taxas. https://t.co/n6mdYYbyiJ
Traduções preparadas para as publicações do portal. Novas publicações e textos alterados aparecem no idioma original até receberem tradução.
Posts sobre IA, modelos e ferramentas, sem respostas nem reposts. Novidades são coletadas duas vezes ao dia; períodos com muitos posts podem levar mais de uma coleta. Imagens, vídeos e conversas completas estão na publicação original. Datas e horários em UTC. Seleções pontuais são identificadas.