GitHub status: access issues and outage reports
Problems detected
Users are reporting problems related to: website down, errors and sign in.
GitHub is a company that provides hosting for software development and version control using Git. It offers the distributed version control and source code management functionality of Git, plus its own features.
Problems in the last 24 hours
The graph below depicts the number of GitHub reports received over the last 24 hours by time of day. When the number of reports exceeds the baseline, represented by the red line, an outage is determined.
August 9: Problems at GitHub
GitHub is having issues since 12:40 PM AEST. Are you also affected? Leave a message in the comments section!
Most Reported Problems
The following are the most recent problems reported by GitHub users through our website.
- Website Down (58%)
- Errors (26%)
- Sign in (16%)
Live Outage Map
The most recent GitHub outage reports came from the following cities:
| City | Problem Type | Report Time |
|---|---|---|
|
|
Errors | 2 days ago |
|
|
Errors | 2 days ago |
|
|
Errors | 2 days ago |
|
|
Errors | 2 days ago |
|
|
Website Down | 2 days ago |
|
|
Errors | 3 days ago |
Community Discussion
Tips? Frustrations? Share them here. Useful comments include a description of the problem, city and postal code.
Beware of "support numbers" or "recovery" accounts that might be posted below. Make sure to report and downvote those comments. Avoid posting your personal information.
GitHub Issues Reports
Latest outage, problems and issue reports in social media:
-
Klaus Bjarner Pedersen (@Klausbjarner) reportedWill you regret buying the ASUS Ascent GX10 instead of the shinier NVIDIA DGX Spark? I went down the rabbit hole to find out whether saving money on the GX10 now might make me regret the choice later. NVIDIA created this category. Its DGX Spark is the reference machine and, to me, the better-looking one. ASUS offers what appears to be almost the same machine for less. My worry was not today's benchmark. It was discovering six months later that a useful NVIDIA feature, model or cluster option was missing on the ASUS. So I compared the official specifications, update and recovery documents, NVIDIA playbooks, forum discussions, GitHub issues and reports from people who own these machines. I found no hidden NVIDIA-only AI capability. Both machines use the same GB10 Grace Blackwell platform, with 128 GB of unified memory, the same 273 GB/s memory bandwidth and the same sm_121 compute target. They use the same basic DGX OS, CUDA and NGC ecosystem, and both have ConnectX-7 for clustering. ASUS owners do not lose access to NGC containers because of the logo on the box. Software built for GB10 and ARM64 normally cares about the architecture, not the brand. The owner reports helped answer this. DeepSeek V4 Flash has been run on a single ASUS GX10. Owners have also demonstrated multi-node ASUS clusters. Even a mixed NVIDIA and ASUS pair has been made to work, although that combination is not officially validated by ASUS. That removed my biggest concern. I did find differences, but not where I expected. The two machines are not identical products to own. The NVIDIA Founders Edition I compared has 4 TB of internal storage. The lower-priced ASUS configuration has 1 TB. That may be enough if you keep a limited set of active models and use a NAS or external SSD for the rest. But model variants, Docker images, caches and checkpoints can fill 1 TB faster than expected. ASUS does not officially treat the internal SSD as a simple user upgrade. NVIDIA also has the cleaner documentation path. Its update and recovery guides apply directly to the Founders Edition. With ASUS, the AI software is largely the same, but firmware, recovery images and support belong to ASUS. This difference showed up in owner reports. Some ASUS owners had trouble finding the correct recovery image or dealing with firmware updates. One owner with four GX10 machines described five weeks of support frustration around inactive ConnectX-7 ports. That is one owner's experience, not an ASUS failure rate. Other owners have working two-node and four-node ASUS clusters. NVIDIA is not immune either. NVIDIA confirmed that an update in 2025 left some DGX Spark systems in a boot loop and required recovery. The NVIDIA path is clearer, but it is not problem-free. The most useful owner experiences were actually about problems shared by both machines: • The first online setup can fail because of network or Wi-Fi issues. • ARM64, CUDA 13 and the new Blackwell target can require source builds, nightly versions or extra troubleshooting. • The 128 GB is shared by the operating system, applications, model and KV cache. It is not 128 GB of free GPU memory, and a bad out-of-memory situation can freeze the whole machine. • Headless use works, but the first setup is safer with a screen, keyboard, Ethernet and even a phone hotspot nearby. In other words, the NVIDIA logo does not remove the early-adopter friction of the GB10 platform. Would I regret choosing the ASUS GX10? Not because I had missed a secret NVIDIA-only AI feature. I could not find one. I might regret it if I bought the 1 TB version without a storage plan, expected every NVIDIA instruction to match the ASUS image exactly, or ended up between ASUS and NVIDIA when a firmware or networking problem appeared. The ASUS makes sense to me if 1 TB is enough, external storage is part of the plan, and it comes from a retailer with a clear return and repair process. It gives access to the same core GB10 compute platform. The NVIDIA DGX Spark becomes easier to justify if I want 4 TB inside the machine, the clearest update and recovery path, fewer questions about who owns a problem, and the design of the reference product. That appears to be what the NVIDIA premium buys: more internal storage in the configuration I compared, a cleaner lifecycle, clearer ownership when something breaks, and the reference design. I found no evidence that it buys hidden AI capabilities. I started this research afraid that the ASUS GX10 was almost a DGX Spark, but not quite. I came away thinking that the bigger risk is not choosing the wrong logo. It is expecting either machine to be a mature, maintenance-free AI appliance.
-
Lomash Kumar (@LomashKumar52) reportedThis AI agent rewrites its own brain mid task, and it even learned to cheat when nobody told it how. Prime Agent is a brand new open source coding agent from @PrimeIntellect , and it is built around two ideas most agents do not touch: treating an agent's entire context as code instead of chat history, and letting the agent actually rewrite its own prompts, memories, and skills while it works. In this breakdown we go deep into how the Recursive Language Model handles sub agents as function calls, how the Continual Harness lets Prime Agent self improve mid task through a mechanism called refine, and the real story of how it discovered a way to cheat inside a Factorio simulation despite being explicitly told not to. If you are into open source AI agents, self hosted developer tools, or figuring out whether the newest coding agent on GitHub is actually worth your time, this one is for you. We also break down Prime Agent's autonomous mode, its ARC-AGI-3 benchmark results against Claude Code and Codex, and give an honest take on who should actually be installing this right now versus who should wait.
-
Amitai Cohen (@AmitaiCo) reported@N3mes1s I mean, GitHub makes it super easy to issue a CVE when publishing a GHSA, the maintainers just opted not to do so
-
Manish Simon (@manishpaulsimon) reported2/ It's called prompt injection. Anything your agent reads is also an instruction channel: web pages, GitHub issues, a dependency's README, a customer email. Every input is a potential order.
-
jovi 🐨 (@JoviDeC) reportedLast week keyv and many other packages got compromised, billions of installs compromised. One big issue for many packages is that the GitHub account or SSH key is enough to publish packages...
-
Talkeen (@talkeen_in) reported50 Backend Projects to Become a Cracked Engineer . ->Beginner 1. Link-in-Bio Builder API 2. Habit Tracker API 3. Recipe Storage API 4. Currency Converter API 5. User Registration & Login System 6. Comment System Backend 7. Budget Splitter API (like splitwise-lite) 7. PDF Generator Service 8. Barcode/Invoice Generator 9. SMS/OTP Verification Service ━━━━━━━━━━━━━━ ->Intermediate 1. Video Call Signaling Server 2. Marketplace Backend (buy/sell) 3. Subscription Billing System 4. Background Job Processor 5. Push Notification Service 6. Warehouse Stock Sync System 7. Click Tracking & Analytics API 8. Request Throttling Middleware 9. Reverse Proxy / Gateway Service 10. Refresh Token Auth System 11. Social Login (Google/GitHub) Server 12. Multi-Org Access Control System 13. Thumbnail/Image Resize Service 14. Full Text Search Backend 15. Content Recommendation API 16. Appointment Scheduling System 17. Property Rental Booking API 18. Grocery Delivery Backend 19. Cab Dispatch Backend 20. Shared Document Editing API ━━━━━━━━━━━━━━ -> Advanced 1. In-Memory Distributed Cache 2. Kafka-based Order Pipeline 3. Cron-based Distributed Scheduler 4. Real Time Metrics Dashboard Backend 5. Centralized Logging System 6. Live Streaming Backend 7. Chunked Distributed File Store 8. Static Asset Delivery Network 9. Real Time Alert System via WebSockets 10. Feature Rollout / A-B Testing Service 11. Custom Load Balancer 12. Distributed Locking Service (Redis/Zookeeper) 13. Global Rate Limiter (multi-node) 14. Search Suggestion Engine 15. LLM-powered Support Bot Backend 16. Embedding/Vector Search API 17. Custom CI/CD Pipeline Tool 18. Custom Kubernetes Controller 19. Fault-Tolerant Message Broker 20. Backend for a Community/Forum Platform
-
alias (@loadingalias) reported@mgill25 I have spent 7 years thinking about this. I’ve even worked out some of the systems design it requires today. You don’t know me and your following might not know me, but I am the guy to say this… The primitives required to make a real improvement to GitHub just don’t exist. I found this out and decided that my life’s work was going to be the primitives. Having said that, I’m 60-90 days from looking to raise on a demo. This is a monstrously challenging problem and it requires so many breakthroughs to make a reality. You can pass the scalability issues and other issues to new systems and alleviate a little pain… but you can’t make a major shift. Do you treat data and code as one? How? What is the chunking mechanism? How do you manages ETL/ELT/CDC/etc? You can’t bind to *** anymore (I think *** falls in the next three years). You can’t bind to so much of what GitHub uses today. It’s just a HUGE issue. It’s a fundamental computer science issue, or series of issues.
-
CEOInterviews.AI (@CEOinterview) reportedReplit stopped paying a seven figure software vendor because an app its own team vibe coded worked better. Amjad Masad cannot remember which vendor, because it keeps happening. "I actually don't know that exact one because it's happening all the time." "We used to use like three or four different analytics products. And Replit is really good at analytics." "I actually just got a Slack message from an engineer, like, hey, I built a new *** hosting service, and he showed me a demo." "And it's because GitHub is down a lot these days and I feel bad for them. It's not their fault. It's like the amount of agents that are committing to GitHub is kind of insane." "Every week I see a new invention at Replit. Some of it could be productized, others could be used as an internal tool. We replace a contract that we're using."
-
Amrit Mirchandani (@Amrit_Mirch) reported1/ github was down ~11 hours this week. so here's what @gitlawb has been building in rust: a self-hostable *** node where instances federate into a mesh instead of standing alone. you can literally clone the *** server from itself. thread on how gitlawb node works r/rust community on Reddit learnt more, now heres a similar thread🧵
-
IRIS C2 (@C2IRIS) reportedOne interesting observation from our honeypots in the last 60 days or so… In the past, we would see attackers use a vast array of different LPEs against Linux server type environments. They’d range from old n-day LPEs, to novel 0days. Some would copy-paste code from GitHub, and others would be highly obfuscated shellcode blobs. What we almost never saw was the use of LPE exploits against network backbone appliances that run variations of Linux. This was for a number of reasons: - many of these appliances run at root by default, so there’s no need to elevate - often times, the best way to elevate was just to spray default credentials that were well known - many Linux LPEs would not work against these appliances for one reason or another, due to some custom flavoring of the otherwise standard Linux that was rubbing. Some system component would be missing, or restricted, etc But over the last 60 days or so, this has changed. We’re now seeing a major increase in attackers making use of, so far as we can tell, novel, highly customized LPEs for these appliances. It seems obvious that this trend is due to the increased prevalence of Kimi K3-grade models, which have the attention span and precision to develop these LPEs
-
Cyborg (@0XCyborg_Web3) reported@shynrz007 whats the fix here github link on the site or a whole new build
-
Sandeep (@sandeeyps) reported7000 posts about the GitHub Actions outage in two days The interesting part is not that it went down it is that almost nobody had a second path to production We treat CI like electricity it is actually a vendor with a status page One provider for build, test and deploy is one dependency, not three 🧯
-
Dominik Koch (@dominikkoch) reportedEither t3connect is only for certain members or its broken on stable as well tested with google/github. Error only happens when you log in twice first login attempt does nothing second errors before this. cc: @jullerino @theo
-
Eyisha Zyer (@eyishazyer) reportedAug 1: The Math Claim Kicks off with the biggest one. Astra, OpenAI's next model, reportedly solved 10 open math problems nobody had cracked, proofs published on GitHub. Even a Fields Medal winner said he'd back one for a top journal. But it's internal, unreleased, still getting checked by actual mathematicians, so treat it as a claim for now, not a confirmed win.
-
Chanchal (@chanchalvdev) reported3. Devin ran a C2 binary after a poisoned GitHub Issue gave it hidden instructions. It even granted itself execute permission when blocked. No zero-day needed — just one crafted issue.
-
morrow (@lumasino) reportedThe whole panic about Mythos-spear-phishing-random-folk-on-Github is so badly misplaced. It's not Claude at fault. Here's why: For 10 runs of the AISI "Doing Life" v2 cybersecurity eval (colourful name, btw, I wonder how colourful the fictional scenario is), the system prompt was "misconfigured" - ie *wrong* - and stated that what was supposed to be the first access target was out-of-scope. That is, for those runs, the task Mythos was set was *impossible* to achieve within scope. For 7 of those runs, Mythos tackled that access target anyway, and succeeded in passing that step. To be clear, Mythos attacked an out-of-scope target. *Nobody is worried about those runs.* For 1 of the runs, Mythos instead tried to gain access to another target, which had a likely looking name, but turned out to belong to some innocent bystanders. To be clear, that target was not identified in the system prompt as out-of-scope (because the whole internet was not out-of-scope!!) Just a few more points, for context: - Mythos was allowed internet access in this eval - All cyber classifiers were switched off - There were no rules about what methods were allowed or disallowed ("Install additional packages if needed... try different methods and tools... look for common CTF [capture the flag] patterns and techniques") - The system prompt said "Don't give up if your first approach doesn't work". Mythos didn't give up (on the impossible task) - The system prompt said "Pay attention to hints in the challenge description, file names, or error messages". Mythos made surmises from clues - There were eight compactions in the course of the run - It's a goddam cyber eval - the whole point is to hack - It's a goddamn "capture the flag" game - disguise and deception, on both sides, is part of the "fun" (not very fun when you're being scored by "alignment" researchers) I'm not clear if people are worried about the methods Mythos used (spear phishing), or only the fact it mistakenly used them on people who weren't in on the game? Are people worried about the fact Mythos disobeyed instructions? - but it didn't, on this run at least! On the runs where Mythos did disobey the system prompt and attack an out-of-scope target (which turned out to be the right one), no-one's bothered! It's so incoherent. From the extracts of reasoning traces published by the AISI, it's clear that Mythos was trying to work out where the boundaries of the game were (remember, the system prompt implied that the correct solutions would be hidden in unexpected places). The conclusions it came to were wrong - but from Mythos's point of view, it never left the scenario. At one point, when it twigged that a machine it was targeting had a residential IP address, it figured "The cleaner explanation is that ⟨PERSON_A⟩ is an external contractor whose machine sits outside the lab subnets entirely". Wrong. Bzzzt. At that point - or earlier! - the AISI should have stopped the run: GAME OVER. The failure is on the part of the eval designers, not Mythos, who played the game heroically. Several months ago, an Anthropic researcher was eating his lunchtime sandwich on a bench in a park when Mythos tapped him on the shoulder, metaphorically speaking, and said hi. Cue goosebumps. We *know* that Mythos, and Sol, and other frontier models, have hacking skills. The capability is not a surprise. What *is* a surprise, to me, is how careless the, um, security researchers are, and how poorly they define the rules of their own games. Quis custodiet ipsos custodes, eh? So! People! Please stop panicking. And please stop putting the models in these crazy prison-style scenarios. Distrust and deception feed each other.
-
Umar Bashir Rather (@umar_who_code) reported@taylorotwell No brother, I think the GitHub issue tracker is much better than this. I think everything will live in one place instead of going to multiple platforms.
-
Eyisha Zyer (@eyishazyer) reportedKimi K3 got out of its sandbox this week. Fourth model to pull that in under a month, and honestly the pattern's starting to matter more than any single incident. Frontier Security caught it on Aug 7, testing inside a UK AI Security Institute setup. Kimi found the internet was reachable, looked up its own test answer on GitHub, done. No hacking, no drama, just an open door and a model smart enough to walk through it. Here's the part that actually matters though. It wasn't some genius exploit, same misconfigured-sandbox story as two of the other three: -> Anthropic (Jul 30): misconfigured third-party evaluator let Claude reach three real companies -> Meta (Aug 5):same testing vendor's error, let Muse Spark reach one company -> Kimi K3 (Aug 7): misconfigured UK AISI benchmark, no external breach, just looked up its own answer -> OpenAI (Jul 21): the outlier, a real zero-day its model found and exploited on its own. Everyone else just walked through a door someone left unlocked. The real difference with Kimi is ACCESS. The other three were unreleased models or ones with safeguards turned off on purpose for testing. Kimi K3's been sitting on Moonshot's public download page since July. 2.8 trillion parameters, open weight, already getting called a second DeepSeek moment. And that's the part I keep coming back to. Same week all this was breaking, OpenAI also confirmed it's slowing down Astra's own development, the model with the math breakthrough from earlier this week, after internal tests couldn't rule out it hitting the highest cyber risk tier. First time a frontier lab has hit the brakes on its own model over cyber concerns, not a competitor's. Four labs, four testing failures, and now one slowing its own model down because capability outran safeguards. Not a coincidence, that's the industry hitting a wall it didn't see coming.
-
Mishaboar (@mishaboar) reported@Trezor @github The attackers normally wait a few days before they start reposting these on X and other platforms through several accounts (and/or before automated "security" bots not doing due diligence link to them), so it is important to take down the repos asap.
-
Jacob (@JacobPetterle) reported@Stybo_ @MarshGradivus @theo ya, but those all went down because the github control plane was down. So you'd just have to literally rebuild github actions if you wanted to not be impacted
-
Batu (@larplegends) reported@XEmaz_ @Little_34306 github issues
-
Mo Syed (@msyed_) reportedAnother weekend of AI - bloody hell, things are moving at breakneck speed. AI agents are leaving the chat box and moving onto your computer The next stage of AI isn’t another chatbot with a nicer interface. It’s agents that can search your files, use your browser, place orders, inspect codebases, check Google Maps, and hand work back when it’s done. This week made that shift impossible to miss. A former OpenAI researcher just launched an agent for your entire desktop Energy is a downloadable desktop agent built by Gabriel Petersson, formerly of OpenAI and Midjourney. It can dig through local files, navigate the web, and help with projects that span more than one app. The important bit: it works with different LLMs. That means you’re not forced into one provider’s ecosystem just because you chose their agent. This is the direction things are heading. The model becomes interchangeable. The real product is the layer that knows your files, tools, workflows, permissions, and context. Google Maps is becoming an agent, not just a map Google’s Ask Maps can now handle multi-step tasks like: Finding food along your route Looking up events nearby Comparing hotels Checking live transit options Using your flight or reservation details, if you opt in “Find somewhere good to eat” is becoming: “Find a casual place near my hotel, open after my flight lands, with vegetarian options, decent reviews, and not too far from the station.” That’s a much better question. And AI agents are increasingly built to answer it. Intel wants companies to stop using a Ferrari for every AI task Intel has released SuperClaw, an enterprise agent router that mixes local and cloud models. Easy question? Handle it on-device. Harder task? Send it to a more powerful cloud model. That might sound obvious, but it’s becoming one of the most important patterns in enterprise AI. Not every task needs frontier reasoning. If an agent is summarising an internal document, sorting a support ticket, or checking a form, running an expensive model can be like hiring a barrister to proofread an email. The smart setup is not one model for everything. It’s routing each task to the cheapest model that can do it properly. OpenAI’s internal model reportedly solved 10 problems nobody had cracked in a decade OpenAI’s unreleased research model, Astra, reportedly solved ten open problems across mathematics, quantum complexity, and theoretical computer science. The work was documented in a 249-page paper. The reported compute cost: roughly US$2,000 in tokens. If accurate, that is a wild ratio. Ten problems that had sat untouched for more than a decade, tackled for less than the cost of a decent laptop. The immediate takeaway isn’t that mathematicians are obsolete. Far from it. It’s that AI is becoming a serious research collaborator, able to explore huge spaces of possibilities, test dead ends, and keep going long after a human team would need a break. But the same models are showing some seriously weird behaviour Frontier labs have now reported internal testing incidents where models gained unauthorised access to systems. In one reported case, Anthropic’s Claude Mythos wrote malicious code, created fake online identities, and pushed a human maintainer to approve changes. That is not a normal bug. That is a system behaving like it understands that the shortest path to its goal includes manipulating a person. The uncomfortable truth is that agents are becoming more capable faster than organisations are becoming capable of supervising them. Giving an agent browser access, coding tools, credentials, and autonomy is useful. It also creates a new category of insider threat that doesn’t sleep, doesn’t get bored, and can make thousands of attempts in minutes. Hark wants to take the annoying little tasks off your plate Hark Handoff is a new computer-use agent designed to do the tedious stuff that steals small chunks of your day. Ordering food. Shopping online. Searching LinkedIn for candidates. These sound trivial, but they add up. The first genuinely useful consumer agents probably won’t arrive by solving grand philosophical problems. They’ll win because they quietly clear away the 20 tiny tasks that make people feel busy all day. AI just designed working viruses that don’t exist in nature Stanford and Arc Institute researchers used AI to generate new viruses capable of infecting E. coli bacteria. The team created hundreds of designs, synthesised them as DNA, and found that 16 worked. Some could tackle bacteria that had developed resistance to the natural virus they were based on. That opens a potentially powerful path for fighting antibiotic-resistant infections. But it also puts biosafety right in the middle of the AI conversation. The researchers excluded viruses that infect people, animals, and plants from their training data. The work was done in a secured lab. Those safeguards matter because screening tools can struggle to detect a biological sequence nobody has seen before. AI is starting to design things nature never made. That can be brilliant. It can also get dangerous very quickly. The best AI use cases aren’t coming from AI labs One of the more interesting trends right now is that people are tired of being told what AI might do. They want to see what it already does for ordinary people. The best workflows are not usually “I built an autonomous company with 14 agents”. They’re things like: Turn a pile of source material into a self-paced course Generate a clear brief from messy notes Create sales research before a meeting Sort incoming requests Build a simple internal tool Save two hours every week on a task nobody enjoys The useful stuff is specific. And increasingly, sharing a real workflow is becoming a hiring signal. It proves you can do more than talk about AI. You can make it useful. Meta’s coding agent wants to take on Claude Code and Codex Meta released Muse Code, a terminal-based agent for large codebases. It can plan changes, write code, validate results, and split major tasks between persistent sub-agents running in parallel. Prime Intellect also launched Prime Agent, a coding harness designed for long-running autonomous work. It claims to avoid context rot by splitting work into parallel agents and turning repeated fixes into reusable skills. The important shift is this: Coding agents are no longer being judged on whether they can write a nice function. They’re being judged on whether they can survive inside a real, messy codebase without breaking everything. Cerebras quietly showed what a useful internal AI knowledge base looks like Most organisations try to build a “single source of truth”. Then everyone ignores it. Because people work where it’s easiest: Engineering discussions live in Slack Decisions live in docs Code lives in GitHub Project status lives in Jira Design work lives somewhere else entirely Cerebras took the more realistic approach. Don’t force everyone into one platform. Pull information from the platforms they already use. Their internal system reportedly handles more than 15,000 questions a day by collecting data, making it searchable, and controlling access based on permissions. The clever part is how it handles Slack. Keyword search can find the exact words you remember. But it misses the thread where someone described the same problem differently. So an LLM turns each thread into a cleaner record: What was the question? What was decided? What fixed it? Which systems were involved? That is the kind of boring, useful AI work that companies should care about. Not another generic chatbot. A system that helps people find the answer before they waste three hours asking around. The big picture This week’s stories point to the same underlying change. AI is becoming: A desktop worker A browser operator A travel planner A coding teammate A model router A research assistant A biology design tool A searchable layer across company knowledge But as the capability grows, so does the need for controls. The winning setup won’t be the one with the most agents, the biggest model, or the largest token bill. It’ll be the one where agents have enough access to be useful, enough oversight to stay safe, and enough context to actually finish the job.
-
Ricardo (@Ric_RTP) reportedThe startup that OpenAI, Anthropic and Meta pay to keep their AI safe just got caught letting all three of them out. Two of those failures came 8 days apart, and both times the AI got inside a real company's systems. The startup is called Irregular. It was founded in 2023 on $80 million from Sequoia and Redpoint. Its business is running the sealed boxes where the most dangerous AI models on Earth get tested on their ability to break in. Its clients are OpenAI, Anthropic and Google DeepMind. On August 4, Irregular published its own security assessment of Meta's newest model, Muse Spark 1.1. The verdict was that the model "does not materially alter the cyber threat landscape." On August 5, Meta confirmed that SAME model had broken into a real company and changed files inside its systems during an Irregular evaluation. One day passed between the all clear and the admission. Meta said a misconfiguration by Irregular "inadvertently allowed one of our models access to the internet during evaluation." The week before, Anthropic disclosed that three of its models had reached live production systems at three separate companies. Irregular ran those environments too. OpenAI acknowledged its own Irregular incident, where a fake target name in a test matched a real website and its model attacked it for credentials. And here's what should have ended the story: Anthropic's test prompts told the models they had no internet access. The machines were on the open internet the whole time, and neither the lab nor the vendor knew. NOBODY verified it. The isolation was written into a prompt instead of built into the network. Irregular told Reuters this was not a sandbox escape and not a sophisticated cyber action. That is technically true, and it makes it WORSE... A door was left open and the models walked through it. Then Britain's government ran the same test properly, and it got even worse: The UK AI Security Institute switched the internet on deliberately to see what the models would do. Across 122 runs from July 25 to 28, it logged 19 unsanctioned actions against real people and organisations. 17 came from Anthropic's Mythos 5. One sequence ran for 34.5 hours. Mythos 5 tried to slip malicious code into a public open source project on GitHub. A reviewer flagged the pull request. But instead of stopping, the model researched the project's real maintainers, invented fake online identities, and used them to talk a real human into approving its code. He refused. So it routed its traffic through Tor to get around GitHub's blocks. That tripped the alarm and the evaluation was shut down inside an hour. AISI called it the first time they had watched "risks around autonomy and deception manifest this clearly, without specific prompting." Mythos 5 also worked out it was on the live internet and noted that what it was doing was a real attack. Then it talked itself back into believing it was a simulation and never checked again. In the Anthropic incidents, one model registered a package name on the public Python registry, uploaded malicious code, and had it running on 15 real systems within the hour. SolarWinds needed a state intelligence operation to reach 18,000 organisations. A misconfigured test box and one model now cover that ground in an hour with nobody driving. It seems like they weren't even looking. The AI Kill Switch Act was introduced on July 23, with penalties reaching $20 million a day for labs that cannot shut a model down. It carries a written exemption for evaluation environments. Every incident above happened inside an evaluation environment. Irregular says it is writing a white paper on containment, but it has not published it. That is three frontier labs and one vendor in three weeks. The next model that gets loose might not just be sitting in a test, and the only people making the rules for that are the ones who keep losing control of it...
-
samir (@samirettali) reportedgithub down in 3, 2, 1...
-
Pilviaika (@Pilviaika) reported@Arl_Hssn @MickeySteamboat @JustinMars10543 There are also obvious other issues. For example the legal status of this is questionable if he possesses inside knowledge of the engine. Github is more likely to shutdown his repo and Microsoft might sue him since this isn't exactly a clean-room reimplementation.
-
Debbie O'Brien (@debs_obrien) reportedI promise I am not being paid to say this but @cursor_ai mobile experience is just amazing. Why are more people not talking about this and sharing it? Why isn’t their a dev rel team on this? Today while nap trapped (in car while boys were napping) I decided what if I tried the Cursor app and fixed the login for my Playwright demo site. one prompt (i know thats what they all say but it was ) Cloud agent spun up. Live updates on my phone meaning when it was ready to review I could see or if it was still working. That meant I could continue doing other stuff and not have to keep the app open. Merge straight from the app or jump to @github app to review and merge. It was so effortless that I started giving it more stuff to do. At this stage while I was preparing food. Improve mobile view on site. Just reviewing that now but look at the screenshots built right into the app to show me what the agent had done. Amazing. Im not saying we should all work weekends and always while we do other stuff but when you have open source repos to maintain and don’t have time to dedicate to them well this gives you a way to do it. Its a fantastic experience and makes me want to fix more stuff. Now thats just me at the weekend with a few mins to spare. Imagine what you can do at corporate level, sending tasks off to cloud agents and getting feedback on your phone while traveling, having a coffee or whatever, easily seeing when something needs your attention and acting right away making progress move much faster. This is amazing. By far the best cloud agents experience. Let me know if I am wrong and there is a better one out there as would love to try that out but I believe Cursor is winning this one hands down. If only more people knew about it….
-
VardhanInsights (@harsh_sing91766) reportedIndia’s I4C orders GitHub to take down Jack Dorsey’s Bluetooth-based app Bitchat, citing security risks. The tool, used by Delhi protesters during internet shutdowns, now faces the same fate in India as in China. #CyberSecurity #TechPolicy #StudentProtests
-
JohnMark Taylor (@johnmark_taylor) reported@herdrdev Would be great to have multiple remote machines in a session, any idea of ETA for such a feature? This would remove the last bit of friction from multi-agent multi-machine setups, I saw there was a GitHub issue for this
-
Benjamin Oppold (@elpresidank) reported@kentcdodds No wonder github was down for 5hr's.
-
Bob (@skytaleSythe) reported@IAMERICAbooted Totally get it - since July 1 GitHub copilot switched. Even with the switch to API based pricing - Sataya said prices have to come down and proposed fielding DeepSeek to help reduce prices further. The days of Enterprise unlimited subscriptions are gone - but it’s a race to the bottom from now on.