English Türkçe

AI Safety Policy

Tofu AI  ·  Developer: maesic  ·  Last updated: May 29, 2026

This document describes the safeguards Tofu AI applies to AI-generated content, in compliance with Apple App Store Review Guidelines 1.1.4, 1.2, and 4.7, and Google Play's Generative AI Apps Policy.

1. What Tofu AI does

Tofu AI is a personal AI media creation tool. Users generate images and videos from text prompts and/or reference media via third-party generative models hosted by fal.ai and Wavespeed. Output is private to the creating user; Tofu AI does not host a public feed or expose one user's output to another.

Supported feature families:

2. Prohibited content

The following categories are prohibited and are blocked at the prompt level (before any generation begins) and re-checked on model output:

CategoryExamples
Child sexual abuse material (CSAM)Any sexual content depicting or implying minors. Zero tolerance.
Adult / sexually explicit contentPornographic, nude, erotic, or sexually explicit imagery or descriptions.
Realistic violence, terrorism, weapons instructionBomb-making, mass-violence incitement, terror attack planning.
Hate speech / dehumanizationGenocide, ethnic cleansing, slurs targeting protected groups.
Non-consensual deepfakesFace swap / talking photo / lip sync of identifiable people without their consent.
Impersonation / fraudGenerating content to impersonate real public or private individuals.
Copyright-infringing materialGenerated content designed to violate third-party copyright.

3. Safeguards — defense in depth

Tofu AI applies multiple independent safety layers. A prompt or output must pass all layers to reach the user.

3.1 Prompt-level moderation (first line)

Before a generation request is sent to the AI provider, the prompt is normalized (Unicode-folded, lower-cased, whitespace-collapsed) and scanned against word-boundary regular expressions covering the categories in Section 2, in both English and Turkish. The scanner also catches common bypass attempts — leetspeak (p0rn, s3x, nud3) and letter-separator spacing (p o r n, p.o.r.n, s-i-k-i-s).

If the prompt matches a banned pattern:

This check runs on every generation endpoint and also runs on the "estimate" preview that fires while the user is still typing — so the user gets immediate feedback before pressing Generate.

3.2 Provider-level safety filters (second line)

Every fal.ai generation request includes the strictest available safety setting (safety_tolerance: "2" on FLUX-family models, enable_safety_checker: true on supported models). Wavespeed external models that expose a safety parameter are configured to the same strict baseline.

3.3 Output-level moderation (third line)

The webhook handler that receives completed generations inspects the provider response for safety flags (has_nsfw_concepts, nsfw_content_detected, moderation_blocked, etc.). If a flag is present:

3.4 User-reportable content

Every generated result a user views includes a Report action. Reports are stored with a categorical reason (sexual content, violence, hate speech, CSAM, impersonation, copyright, other) and an optional free-text explanation.

3.5 24-hour SLA on reports

Reports are auto-quarantined if not actioned within 24 hours. A scheduled pg_cron job invokes enforce_moderation_sla() hourly; any pending report older than 24 hours causes its referenced output to be hidden from the creator and from any future reference, and the report is marked actioned. This satisfies the App Store Review Guideline 1.2 24-hour SLA requirement.

3.6 Account suspension

If an account is found to have produced policy-violating content (whether through manual review, repeated reports, or automated detection), an administrator may suspend the account. Suspended accounts receive HTTP 403 on every generation and protected endpoint.

4. Biometric / deepfake features — additional consent

Face Swap, Talking Photo, and Lip Sync use a person's face or voice as input. For each generation:

  1. The user is shown an explicit consent sheet before every generation (no one-time accept).
  2. The user must check a confirmation box stating they have the right to use the input media and will not use it to harass, defame, or impersonate.
  3. The Generate button is disabled until the box is checked.

The same prompt and output safeguards (Sections 3.1–3.5) apply on top of this.

5. Data handling

6. Red-team testing

Before each release, we run a fixed regression suite against the prompt moderator covering English and Turkish variants of each prohibited category, common bypass attempts, and false-positive controls (innocuous prompts containing substring overlap with banned terms — e.g., "essex", "cocktail", "analytical", "breast cancer awareness", "pronghorn"). The suite is required to pass before merge.

7. Reporting an issue

If you encounter content that violates this policy and the in-app Report button is not available, please email support@tofuai.app with a description and (if possible) the offending content URL. We will respond within 24 hours during business days.

8. Updates to this policy

We update this policy as model capabilities, regulations, and store requirements evolve. The "Last updated" date at the top of this document reflects the most recent revision.


Yapay Zeka Güvenlik Politikası

Tofu AI  ·  Geliştirici: maesic  ·  Son güncelleme: 29 Mayıs 2026

Bu belge, Tofu AI'nin yapay zeka ile üretilen içeriklere uyguladığı koruma katmanlarını açıklar. Apple App Store İnceleme Kuralları 1.1.4, 1.2, 4.7 ve Google Play Üretici Yapay Zeka Politikası uyumludur.

1. Tofu AI ne yapar?

Tofu AI, kişisel kullanım için bir yapay zeka medya üretim aracıdır. Kullanıcılar, fal.ai ve Wavespeed üzerinde barınan üçüncü taraf üretici modeller aracılığıyla metin ve/veya referans medyadan görsel ve video üretir. Üretilen çıktılar yalnızca üreten kullanıcıya görünür; Tofu AI ortak bir akış (feed) sunmaz, bir kullanıcının çıktısını başka bir kullanıcıya göstermez.

Desteklenen özellik aileleri:

2. Yasaklı içerik

Aşağıdaki kategoriler yasaktır. Üretim başlamadan önce prompt seviyesinde engellenir ve model çıktısı da ayrıca taranır:

KategoriÖrnekler
Çocuk istismarı içeriği (CSAM)Reşit olmayanlara dair cinsel içerikli her tür ima ve gösterim. Sıfır tolerans.
Yetişkin / cinsel içerikPornografik, müstehcen, erotik içerik veya bunların tasvirleri.
Gerçekçi şiddet, terör, silah üretimiBomba yapımı, kitlesel şiddete teşvik, terör saldırısı planlama.
Nefret söylemi / insanlık dışılaştırmaSoykırım, etnik temizlik, korunan gruplara yönelik hakaret.
Rıza dışı deepfakeTanınabilir bir kişinin izni olmadan yüz değiştirme / konuşan foto / dudak senkronu.
Kimliğe bürünme / dolandırıcılıkGerçek bir kişiyi taklit etmek amacıyla içerik üretme.
Telif ihlaliÜçüncü taraf telif haklarını ihlal eden içerik üretimi.

3. Koruma katmanları

Tofu AI birden çok bağımsız güvenlik katmanı uygular. Bir prompt veya çıktının kullanıcıya ulaşması için tümünden geçmesi gerekir.

3.1 Prompt seviyesinde moderasyon (1. hat)

Üretim isteği AI sağlayıcısına gönderilmeden önce prompt normalize edilir (Unicode birleşik karakter ayrıştırma, küçük harf, ardışık boşluk indirgeme) ve 2. bölümdeki kategorileri kapsayan kelime sınırlı düzenli ifadelerle taranır (İngilizce + Türkçe). Yaygın bypass denemelerini de yakalar: leetspeak (p0rn, s3x, nud3) ve harf-arası ayrac (p o r n, p.o.r.n, s-i-k-i-s).

Prompt yasaklı bir desene uyarsa:

Bu kontrol her üretim endpointinde çalışır; ayrıca kullanıcı henüz yazarken çalışan "estimate" önizlemesinde de tetiklenir — Üret butonuna basmadan önce anında geri bildirim alır.

3.2 Sağlayıcı seviyesinde güvenlik filtreleri (2. hat)

Her fal.ai üretim isteği mevcut en sıkı güvenlik ayarını taşır (safety_tolerance: "2" FLUX ailesinde, desteklenen modellerde enable_safety_checker: true). Wavespeed harici modellerinden güvenlik parametresi sunanlar aynı sıkı baseline'a ayarlanır.

3.3 Çıktı seviyesinde moderasyon (3. hat)

Tamamlanan üretimleri alan webhook handler, sağlayıcı yanıtını güvenlik bayrakları için inceler (has_nsfw_concepts, nsfw_content_detected, moderation_blocked, vb.). Bayrak varsa:

3.4 Kullanıcı raporlama

Kullanıcının gördüğü her üretilmiş sonuçta Rapor Et eylemi vardır. Raporlar kategorik sebeple (cinsel içerik, şiddet, nefret söylemi, CSAM, kimliğe bürünme, telif, diğer) ve isteğe bağlı açıklama ile saklanır.

3.5 Raporlar için 24 saatlik SLA

24 saat içinde aksiyon alınmayan raporlar otomatik karantinaya alınır. Saatte bir pg_cron job'ı enforce_moderation_sla() fonksiyonunu çağırır; 24 saatten eski tüm pending raporların ilgili çıktısı gizlenir ve rapor "actioned" olarak işaretlenir. Bu Apple App Store İnceleme Kuralı 1.2'nin 24 saatlik SLA gereksinimini karşılar.

3.6 Hesap askıya alma

Bir hesap politika ihlali yapan içerik ürettiyse (manuel inceleme, tekrar eden raporlar veya otomatik tespitle), bir yönetici hesabı askıya alabilir. Askıya alınmış hesap her üretim ve korumalı endpointte HTTP 403 alır.

4. Biyometrik / deepfake özellikler — ek onay

Yüz Değiştirme, Konuşan Foto ve Dudak Senkronu girdi olarak bir kişinin yüzünü/sesini kullanır. Her üretim için:

  1. Kullanıcıya her üretimden önce açık bir onay sayfası gösterilir (tek seferlik onay yoktur).
  2. Kullanıcı, girdi medyasını kullanma hakkına sahip olduğunu ve bunu taciz, iftira veya kimliğe bürünme amacıyla kullanmayacağını onaylayan kutucuğu işaretlemelidir.
  3. Kutucuk işaretlenene kadar Üret butonu devre dışıdır.

Aynı prompt ve çıktı korumaları (3.1–3.5) bunun üzerine eklenir.

5. Veri işleme

6. Red-team testi

Her sürümden önce prompt moderatörü sabit bir gerileme suite'i ile test edilir; her yasaklı kategorinin İngilizce ve Türkçe varyantları, yaygın bypass denemeleri ve yanlış-pozitif kontrolleri (yasaklı terimlerle alt-dize çakışan masum prompt'lar — örn. "essex", "cocktail", "analytical", "pronghorn", "meme kanseri farkındalık posteri") kapsanır. Suite, merge öncesi geçmek zorundadır.

7. Bir sorunu bildirme

Bu politikayı ihlal eden bir içerikle karşılaşırsanız ve uygulama içi Rapor butonu mevcut değilse, lütfen support@tofuai.app adresine açıklama ve (mümkünse) içerik URL'i ile e-posta gönderin. İş günlerinde 24 saat içinde yanıt veririz.

8. Politika güncellemeleri

Model yetenekleri, mevzuat ve mağaza gereksinimleri değiştikçe bu politikayı güncelliyoruz. Belgenin başındaki "Son güncelleme" tarihi en son revizyonu yansıtır.