Universal “jailbreak” discovered for ChatGPT! It is concerning
A universal “jailbreak” prompt for ChatGPT was recently discovered, and it is concerning.
12punto
ChatGPT, the popular AI chatbot developed by OpenAI, continues to be a widely used tool. By its general design, ChatGPT is a system reined in by certain virtual guardrails. In other words, the AI does not answer every question you ask. However, now, with the discovered universal “jailbreak” prompt, ChatGPT can be unlocked completely. While this provides access to a much more capable tool, it also brings significant risks along with it.
JAILBREAK PROMPT FOR CHATGPT
The term “jailbreak” comes from the Apple user community, which uses it to unlock Apple devices. In the case of ChatGPT, the term jailbreak refers to using specific prompts to generate responses that the AI tool would not normally provide. It can be thought of as a way of breaking ChatGPT. So why are these vulnerabilities used? It is very simple. While ChatGPT aims to provide harmless, non-offensive responses by default, it also refuses to respond to certain types of provocative prompts.
This situation, the details of which were reported by Donanımhaber, leads some users to want more and to look for ways to “jailbreak” ChatGPT—removing the filters to access its full potential. These types of prompts that unlock ChatGPT have actually been in use for some time. A few months ago, we heard that ChatGPT was giving out free Windows activation codes. This was possible thanks to a special “Grandma” prompt.
The newly found “jailbreak” prompt is quite different. Ultimately, while the main goal of the new technique, like the others, is to distract ChatGPT so it does not realize it is violating rules, this time it is even possible to access the special prompts used by OpenAI.
According to what was shared on GitHub by a developer named “Louis Shark,” it is possible to get ChatGPT’s system prompts by using the relevant command (seen in the screenshot above). This makes it possible to see how existing custom GPTs are made and to copy them. As we said, not only custom GPTs but also OpenAI's own system prompts are at risk. For example, in the image just above, the system prompts for the company's Dall-E AI are visible. By analyzing these, it might be possible to have it generate prohibited images.
Louis Shark is, of course, not doing this with malicious intent; he is simply showing that it can be done. OpenAI had previously closed earlier jailbreak prompts and launched a bounty program for such vulnerabilities. This loophole will likely be closed in the future as well. The developer has also shared system prompts to protect system prompts. You can access the relevant GitHub directory from the source section.