30 Eylül 2026 Çarşamba

JSON Çeşitleri

Giriş
Açıklaması şöyle
There's a whole ecosystem of JSON-adjacent formats that have emerged over the years, each extending or modifying the base specification to solve problems the original format wasn't designed for.
JSONL  - JSON Lines
Açıklaması şöyle. Streaming ve memory problemini çözer
JSONL is multiple independent JSON objects, one per line. Each line is a complete, valid JSON record. A thousand records means a thousand lines, each standing alone.
Örnek
Şöyle yaparız
// Regular JSON - one structure
{
  "records": [
    {"id": 1, "name": "Alice", "status": "active"},
    {"id": 2, "name": "Bob", "status": "inactive"},
    {"id": 3, "name": "Carol", "status": "active"}
  ]
}
// JSONL - one record per line
{"id": 1, "name": "Alice", "status": "active"}
{"id": 2, "name": "Bob", "status": "inactive"}
{"id": 3, "name": "Carol", "status": "active"}

JSON5 ve JSONC - JSON with Comments
Açıklaması şöyle
Allows comments and trailing commas for more forgiving configuration files. JSONC (JSON with comments, used by VSCode) does something similar.
BISON - Binary JSON
Açıklaması şöyle
Powers MongoDB's internals.
EJSON (Extended JSON)
Açıklaması şöyle
Adds additional type support. 
GeoJSON
Açıklaması şöyle
Structures geographic data.



28 Eylül 2026 Pazartesi

Yazılım Mimarisi - Replication Consistency

Giriş
Açıklaması şöyle
The primary replica set up will result in update delay in replicas and is a classic eventual consistency model. Essentially we trade strong consistency for read scalability. Eventual consistency is enough for most applications, except for ones requiring ‘read your write’ consistency.

‘Read your write’ consistency can be improved by forcing the read request to primary if it’s following a write. Or naively force the read to wait for several seconds so that all replicas have caught up. When there are replicas not in the same datacenter(DC), the read will also need to be restricted to the same DC.
Read-your-own-writes consistency 
Farklı çözümler var

1. Route To Master
Açıklaması şöyle
... 
a user updates their profile and doesn't see the change.
...
But there's a catch nobody warns you about: replication lag.

Your replica is 200ms behind master. User changes their avatar, page reloads, reads from replica - old avatar. "Did my update even save?" They click save again. Now you have a duplicate write and a confused user.

The fix is called read-your-own-writes consistency. After a write, route that specific user's reads to master for the next N seconds. Everyone else still reads from replicas. This solves 90% of "my changes disappeared" tickets.

Implementation: set a short-lived cookie or Redis key after a write. Middleware checks it - if present, route to master. If expired, back to replica. Five lines of code that save you hundreds of bug reports.

2. Short TTL Redis Kullanmak
Örnek
Şöyle yaparız
// CORRECT — a user's own recent writes are read from the primary
@Service
@RequiredArgsConstructor
public class ProfileService {
    private final RecentWriteTracker recentWrites; // Redis, short TTL
    @Transactional
    public void updateProfile(Long userId, ProfileRequest request) {
        User user = userRepository.findById(userId).orElseThrow();
        user.applyChanges(request);
        userRepository.save(user);
        recentWrites.mark(userId, Duration.ofSeconds(10)); // Longer than max lag
    }
    public UserProfile getProfile(Long userId) {
        return recentWrites.hasRecentWrite(userId)
            ? readFromPrimary(userId)   // This user just wrote - don't risk the replica
            : readFromReplica(userId);  // Everyone else reads the cheap path
    }
}


25 Eylül 2026 Cuma

Redis - Bitset Veri Yapısı

Giriş
Komutlar şöyle
SETBIT
GETBIT
BITCOUNT
BITOP
BITPOS

SETBIT
Bir kullanıcının yılın hangi gününde giriş yapıtğını atamak için şöyle yaparız
SETBIT sign:123 10 1

Burada sign:123 key yani kullanıcı adı, 10 ise bitoffset ve bit değeri 1

GETBIT
Okumak için şöyle yaparız
GETBIT sign:123 10

BITCOUNT
Bu kullanıcı yılın kaç günü giriş yapmış görmek için şöyle yaparız
Okumak için şöyle yaparız

BITOP
İki tane kullanıcının bitsetlerini karşılaştırmak için şöyle yaparız
BITOP AND result sign:d123 sign:124

sonra şöyle yaparız
BITCOUNT result

23 Eylül 2026 Çarşamba

Cache Stratejileri - Cache Stampede

Giriş
Açıklamalar çoğunlukla buradan.

1. No protection
Every caller discovers the miss and independently loads the value.

This is the classic cache stampede / thundering herd.

The dangerous part is that the application may be completely correct. The database is simply receiving an unexpected burst.

2. Local singleflight
Here, each application instance says:

"If another thread on this same machine is already loading this key, I'll wait for its result."

Burada halen pod sayısı kadar hit gelebilir.

3. Fleet refresh lease
Only the pod that acquires the lease is allowed to refresh. Burada A uyandı ancak işi bitiremedi, sonra B uyandı işi bitirdi ve daha sonra A uyandı o da işi bitirdi ancak eski veriyi yazdı problemi var. Yani optimistic lock koymak lazım

4. Stale-while-revalidate
you distinguish:
- fresh
- stale-but-usable
- too-old / unavailable

Stale iken de veriyi sunar

The TTL jitter
Hepsi için kullanılabilir. Açıklaması şöyle
Bad jitter
If you do:
TTL = maxTTL + random()

then everything is already guaranteed to live at least maxTTL.

You're spreading the expiration times, but you're also violating the intended freshness boundary.

Subtractive jitter
Instead:
TTL = maxTTL - random()

gives:
maxTTL-0
maxTTL-3
maxTTL-17
maxTTL-42
...

So everything expires at or before the maximum freshness deadline.

The important concept isn't really the exact formula.

It's:

Randomize expiration without extending the maximum allowed lifetime.







18 Eylül 2026 Cuma

Workflow vs Use case

Workflow
The workflow is the long-running process required to accomplish an operation
Use case
A use case represents a business action.
Şöyle düşünebiliriz.
Use case:
    create/update/delete business object
    apply business rule
    complete a business operation
    produce business events
Book Keepin
The retry count, lease, sent flag, processing flag, backoff timestamp, etc. exist solely so the workflow can safely coordinate work.

Örnek
Use case : Register User

Implementation:

- Validate email is unique
- Create user
- Assign default role
- Store user
- Publish UserRegistered event
Workflow şöyle. The workflow spans time, retries, external systems, and multiple transactions.
User Registration Workflow

  - Register User use case
  - Send verification email
  - Wait for user click
  - Timeout after 7 days
  -  Resend email
  -  Activate account
Db Book Keeping Operations
Hem workflow, hem de use case içinde book keeping transactions olabilir. Bunlar use case anlamına gelmezler.

Örnek
X Delivery Workflow
│
├─ find pending delivery
├─ mark SENDING                 ← DB operation
├─ send UDP segment 1
├─ wait for ACK
├─ send UDP segment 2
├─ wait for ACK
├─ timeout?
│    └─ mark BACKOFF            ← DB operation
│
├─ retry
│
├─ final ACK
│    └─ CompleteDelivery    ← Use case
│
└─ publish resulting events


9 Eylül 2026 Çarşamba

Inventory Reservation Checkout

Giriş
Açıklaması şöyle
Shopify replaced Redis with MySQL for one of their most critical checkout workloads. The trick was surprisingly simple.


Every time a buyer clicks "Complete purchase," Shopify has milliseconds to answer one question: is that last unit still available?


If Shopify says yes when the item is already gone, two people could buy the same item, leading to a cancelled order, an apology email, and a loss of trust.


If they say no when the item is actually available, they lose a genuine sale.


For years, Shopify used Redis to handle inventory reservations during checkout. When a customer started buying an item, Shopify would reserve that item in Redis so another customer couldn't buy the same unit at the same time.


But the actual inventory lived in MySQL, so reservations and inventory updates happened in two different systems with no shared transaction. A payment could succeed while the inventory update failed, causing an oversell, or the inventory could be deducted while the Redis reservation remained, causing an undersell.


The obvious fix was to move reservations into MySQL too, so everything could happen in one atomic transaction.


But MySQL had already failed at this once. Their first attempt used a single row with a quantity column, and at checkout scale, every buyer fought over that one row. Locking it, waiting, retrying. It buckled.


This is where the magic comes in.


Instead of using one row per item with its available quantity, they used one row for each individual unit. So an item with 10 units gets 10 rows. To reserve three units, they simply grab any three available rows.

And then the key: MySQL 8's 𝗦𝗞𝗜𝗣 𝗟𝗢𝗖𝗞𝗘𝗗.


Normally, if a query wants rows another transaction has locked, it waits, everyone queues behind the same rows, and throughput collapses. SKIP LOCKED flips that: if some rows are locked, MySQL simply skips them and hands back other free ones. No waiting. No queue on a hot row.


So when 500 buyers check out the same product at once, they're not fighting over one row, they each quietly grab a different free unit and move on. Contention, which killed the naive design, basically disappears.


To keep it fast at huge scale, they cap the pool at ~1,000 rows per item and refill it from the ledger as it drains. Same idea, bounded.


That's the whole trick: turn "everyone fights over one number" into "everyone grabs a different row." A single database feature made MySQL viable for a workload people assumed needed specialized infrastructure.


The boring database you already run is often more capable than you think. MySQL didn’t just handle Shopify’s scale, it also made their system much simpler.

Caddy Reverse Proxy

Change host file
C:\Windows\System32\drivers\etc\hosts
Add

#Caddy
127.0.0.1 foo.com.local

Start caddy
caddy fmt --overwrite --config .\Caddyfile
caddy run --config .\Caddyfile

Import Root CA
caddy trust

To check Root CA go to
C:\Users\acelya\AppData\Roaming\Caddy\pki\authorities\local
You will see root.crt and root.key files

My certificate is at
C:\Users\acelya\AppData\Roaming\Caddy\certificates\local\foo.com.local


Check Root CA
Run
certmgr.msc

Go to "Trusted Root Certification Authorities -> Certificates"
You should see Caddy Local Authority

Visit
https://foo.com.local

Import Root CA To Remote Clients
Copy root.crt file to other client computers and install it into the "Trusted Root Certification Authorities" store
Command to import on Windows (run as Administrator)
certutil -addstore -f "ROOT" root.crt

Or via GUI
- Double click root.crt -> Install certificate
- Select Local Machine -> Place all certificates in the following store
- Browse to Trusted Root Certification Authorities -> OK -> Finish

Why root.crt and not intermediate.crt ?
- root.crt - The root CA certificate. Clients trust this once, then autimatically trust every certificate Caddy issues
- intermediate.crt - signed by root.crt, but clients need the root in their trust store to validate the chain. Distributing the intermediate alone doesn't help unless the root in already trusted.

Örnek Caddyfile
# Caddy auto_https normally binds to port 80 for HTTP -> HTTPS redirects
# disable_redirects keeps TLS on 443 but prevents binding to port 80
# Needed on Windows where port 80 required admin privileges
{
  auto_https disable_redirects
}

# HTTPS-only listener on port 44
foo.com.local:443 {
  tls internal

  # Swagger UI-served by the Spring Boot app
  handle /swagger-ui/* {
    reverse_proxy localhost:8080
  }

 # OpenAPI docs served by the Spring Boot app
  handle /v3/api-docs/* {
    reverse_proxy localhost:8080
  }
 # REST endpoints
  handle /api/* {
    reverse_proxy localhost:8080
  }
 # Everything else -> frontend
  handle {
    reverse_proxy localhost:3000
  }
}



13 Ağustos 2026 Perşembe

Cuckoo Filter -

Giriş
Bloom Filtreden farklı olarak silmeyi destekler. Açıklaması şöyle
A Cuckoo filter is a probabilistic data structure that supports approximate set membership with the ability to add, lookup, and delete items. It stores a small fingerprint of each item in a fixed-size bucket array. Lookups tell you:

- Definitely not present — you can trust this and skip the expensive check.
- Possibly present — you must verify against the source of truth.

It’s called “cuckoo” because it uses cuckoo hashing: each item has two candidate buckets. If both are full during insertion, the filter evicts an existing fingerprint and moves it to its alternate bucket, potentially triggering a cascade of evictions — like a cuckoo bird pushing eggs out of a nest.


10 Temmuz 2026 Cuma

Positional Block Encoding - Hierarchical Block Allocation

Giriş
Verilen n kademe/seviye girdiyi tek bir sayıya çevirir

1. İki Kademe
Formül şöyle
- areaNo
- systemNo
- SYTEMS_PER_AREA - Yani base
- TERMINALS_PER_SYSTEM
- BLOCK_SIZE

Start : (areaNo * SYTEMS_PER_AREA + systemNo) * TERMINALS_PER_SYSTEM
End : Start + BLOCK_SIZE
Yeni bir sabit hesaplarız
TERMINALS_PER_AREA = SYTEMS_PER_AREA * TERMINALS_PER_SYSTEM
Sonra
start = (areaNo * TERMINALS_PER_AREA) + (systemNo * TERMINALS_PER_SYSTEM)
ve end
start + BLOCK_SIZE


1 Haziran 2026 Pazartesi

DMR (Digital Mobile Radio) Notlarım

Kişisel Notlar

DMR standardı, ETSI, European Telecommunications Standards Institute tarafından yazılmış.

İngilizce kullanımda portable radio el telsizi demek. Yani

Portable Radio - El telsizi
Mobile Radio - Araç Telsizi

Tiers
Açıklaması şöyle
Understanding the DMR Tiers is not strictly necessary but knowing which Tier your radio is, is very important.

TIER 1 - Radio to Radio
  Simplex only. No Time slots. This means it cannot be used on amateur repeaters as it would use up both time slots.
The original Boefeng DM5R was Tier 1 only. DM5R+ had Tier 2.
TIER 2 - Repeater-based
   Supports 2 Time Slots (TDMA). What us Amateurs Use
TIER 3
Advanced Trunked system. Complicated Commercial system allows for automatic repeater switching plus many other things.
Similar to the Airwave system used by the emergency services.
Tier 2 - Conventional DMR  - Single Site
ETSI TS 102 361-2 dokümanında Tier II conventional air interface açıklanıyor. Aslında repeater centric olarak düşünülebilir.
Site yoktur. Şeklen şöyle
Talkgroup <-->  Repeater
Tier 2 - Wide-Area DMR - Multi Site - Vendor Specific
Site vardır. Siteler, bunların tanımları, nasıl roaming yapıldığı vs. hep üreticiye özel.
Sitedeki repeaterlar IP ile (e.g. IP Site Connect) birbirlerine bağlanır.
Şeklen şöyle. Burada telsizler sitelerle ilişkilenir. Sitelerin kimlik numarası vardır ve birbirlerine yönlendirmeyi ağ yönetim Ip uygulaması bilir ve yapar. Yani routing statiktir.
Talkgroup <-->  Site <--> Repeaters
Tier III - Trunk DMR
ETSI TS 102 361-4 dokümanında  açıklanıyor. Site, Control channel, Traffic channels, Base stations (repeaters), Site identifiers, Registration, Roaming between sites,  Dynamic channel/timeslot assignment gibi şeyler açıklanıyor

Channel
Kanala bazı özellikler atanabilir. Bunlar üreticiye mahsus şeylerdir
- Kanal seçilince otomatik olarak bir Talk Group dinlenmeye başlanır
- Kanal seçilince otomatik olarak kriptolu/açık iletişim kullanılır

Channel Plan
Açıklaması şöyle
A channel plan is the mapping between logical channel numbers (LCNs) and actual RF frequencies used by a DMR system.
Açıklaması şöyle
The DMR standard defines the concept of logical channels and how channel information can be communicated, but the actual channel plan (the specific frequencies and LCN assignments) is determined by the system operator or vendor deployment.
Açıklaması şöyle
A grant such as:

> Talkgroup 1001 → LCN 2 → Slot 1

means:

> Tune to 451.0125 MHz and use Slot 1.
 OTAP — Over-The-Air Programming
Açıklaması şöyle
OTAP allows a radio's configuration (codeplug) to be updated remotely through the radio network instead of connecting a programming cable.
OTAP protokolü üreticiye mahsustur ve muhtemelen all radios gönderim yeteneği vardır.

OTAR — Over-The-Air Rekeying
Açıklaması şöyle
OTAR is completely different. OTAR updates cryptographic keys, not radio configuration.
CSBK - Control Signalling Block
Bir CSBK mesajı şu hedeflere adreslenebilir
A specific radio (Individual ID)
A talkgroup (Group ID)
All radios (Broadcast)
A subset of radios depending on the CSBK type
DMR 24-bit adresler kullanır. Bazı özel adresler şöyle
0xFFFFFF (16777215) All Radios / Broadcast / All Call
0xFFFFFE (16777214) Reserved special address
0x000000 Often used as Null Address depending on context
CSBK Mesaj Grupları
The CSBK contains a 6-bit CSBKO (CSBK Opcode) field. The opcode determines the message family and how the 64-bit data field is interpreted.

Common ETSI CSBK families include:

Group                                         Purpose
Channel Grants                         Assign traffic channels for voice or data
Announcement Messages         Broadcast system information
Voting / Adjacent Site Messages Roaming and site selection
Random Access / ALOHA         Registration and access control
Acknowledgement Messages         Confirm requests
Negative Acknowledgement           Messages Reject requests
Maintenance / Control                 System management
Manufacturer Specific                 Vendor extensions using MFID
OpCode dışında FID alanı da önemli. Açıklaması şöyle
Also note that ETSI defines a second level of grouping through the FID (Feature ID) field. The same CSBKO value can have different meanings depending on the FID. ETSI reserves FIDs for standard features and manufacturer-specific features.
Simulcast Zone
Açıklaması şöyle.
Simulcast zone, birden fazla fiziksel siteyi tek bir RF hücre (logical site/cell) gibi çalıştıran senkronize yayın domain’idir.
Şeklen şöyle. Bu yüzden siteler arasında roaming olmuyor.
Site A ─|
Site B ─|── same frequency, same slot, same content
Site C ─|
Açıklaması şöyle. Yani tekrarlayıcıların (repeater) senkronizasyonu çok önemli.
The challenge with simulcast is that a radio may receive signals from A, B, and C simultaneously. Therefore the transmitters must be synchronized very precisely (typically GPS-disciplined timing and frequency references). Otherwise the overlapping signals create destructive interference.
Talk Group
Açıklaması şöyle
A DMR GROUP  (Often called a “talkgroup” ) is a method of grouping or assembling multiple users (Radio ID’s) to a single contact. A Group or “talk group” is simply a group of users that need to talk to each other and hear all the communications in that group. 

- An example would be all the maintenance personnel at a hospital would need to all share communications with each other. So, you would create the  “maintenance” group.
Broadcast: Bazı gruplar abone üye olsun olmasın global kabul edilir ve tüm telsizlere yayın yapar
Roaming : Telsiz site değiştirdikçe roaming bayrağına göre bu gurubu dinlemeye devam eder. 

Tier 1 – RF Control & Channel Architecture (DMR Systems)


1. Trunking Architecture Types

1.1 Site-Controlled Channel Allocation

  • Each site manages its own channel allocation
  • Decisions are made locally at site level
  • Traffic channel assignment is handled by the site
  • No mandatory centralized routing entity

1.2 Centralized Architecture (Star Topology)

Central system is managed by ISS (Integrated Soft Switch) in a star topology.

Topology
  • Star topology
  • All sites connect to central ISS
  • Routing decisions are centralized
Call / Channel Allocation Flow
  1. Radio locks to site control channel
  2. Site forwards registration to ISS
  3. ISS performs authorization and routing decision
  4. ISS instructs site to assign traffic channel
  5. Site executes RF channel assignment

Summary: Site executes, ISS decides.


2. Channel Types

2.1 RF carrier bandwidth

Dijital kanal 12,5 kHz bant genişliğindedir. Analog kanal DMR standardı değildir ama kullanılır ve genellikle 25 kHz bant genişliğindedir

2.2 Analog Channel

  • Not part of DMR traffic
  • Supported for legacy systems

2.3 DMR Conventional Channel

Simplex (P2P)

  • Radio-to-radio communication
  • No repeater
  • Direct communication model

Repeater Mode

  • Uses repeater infrastructure
  • Extends coverage
  • Frequency pair operation

2.4 DMR Trunked Channel (Tier III)

  • Dynamic channel allocation
  • Separation of control and traffic planes
Channel Types:
  • Control Channel: signaling and system control
  • Traffic Channel: voice/data communication

3. Squelch (Muting Mechanism)

Squelch is a receiver function that mutes speaker output when no valid signal is present. Without it, constant background noise ("hiss") would be heard.


4. Squelch Types

4.1 Color Code (Digital Squelch)

  • Used in digital DMR channels
  • Equivalent to CTCSS/DCS in analog systems
  • Range: 1–15
  • Filters unrelated digital systems on same frequency

4.2 Tone Squelch (CTCSS – Analog)

  • Uses sub-audible tones (e.g. 88.5 Hz)
  • Must match for audio to open
  • Allows frequency sharing without interference

4.3 Digital Squelch (DCS – Analog)

  • Digital code-based squelch
  • Alternative to CTCSS
  • Uses digital patterns instead of tones

5. ANI - Automatic Number Identification

Sadece analog kanallarda olur. Analog tone ile arama yapan telsiz kimliği karşı tarafta gösterilir.

5. SelectiveCall

Sadece analog kanallarda olur. Analog tone ile aranan telsizin ses çıkarması sağlanır.

6. Kripto

Aslında DMR standardında kripto tanımlı değil. Çoğu üretici kanala bağlı bir kripto tanımlıyor. Telsiz kanal değiştirince, o kanal için tanımlı kripto algoritmasını kullanıyor

7. Repeater

Repeater'lar katmanlı çalışır ve birbirlerine trafik gönderebilirler

8. System Mental Model

More to come

21 Nisan 2026 Salı

Monitoring vs. Observability

Giriş
Açıklaması şöyle
Monitoring answers: "Is something wrong?" and observability answers: "Why is it wrong?" You need both.
Açıklaması şöyle
In practice, you need three things. Metrics are used to detect problems, logs to explain errors, and traces to discover latency. If your metrics are wrong, you would never know that something is failing. And if you don't know something is failing, you never check logs and traces, which is why metrics are the entry point of any investigation.
Monitoring Challenges
Açıklaması şöyle
Most of the time, teams don't have strategies for monitoring. It is the last backlog item to be picked up before the final production release. One service team adds a dashboard, another adds alerts, and a third team introduces a different naming convention.

Six months down the line, you get duplicate metrics and inconsistent naming. There are no standard dashboards and alerts that nobody trusts. Eventually, teams ignore alerts, stop relying on monitoring, and fall back to guesswork. That is a dangerous place to be.

One pattern I have seen repeatedly is metric explosion without clarity. A service exposes 400 metrics, and nobody knows which one matters.

Good monitoring is not about collecting more metrics. It is about collecting the right metrics. A production-ready service rarely needs more than 10-20 core metrics and a small number of critical alerts. Everything else is an investigation detail. Not an operational signal.
Yazar 4 tane metriğin takip edilmesi gerektiğini söylüyor
1. Latency: Earliest Signal
Burada ortalamaları (average) gösteren metriklerin işe yaramadığını söylüyor. Percentile p50, p90 gibi bakmak daha iyi. Eğer p99 hareket etmeye başlarsa bir problem var anlamına gelir.
Örnek
Şöyle yaparız. order-service içindeki end pointleri gösterir.
uri : groups metrics per endpoint, 
le : “less than or equal”
anlamına gelir.
# Latency (p50, p95, p99)
histogram_quantile(0.50, 
  sum(rate(http_server_requests_seconds_bucket{application="order-service"}[5m])) by (uri, le)
)
histogram_quantile(0.95, 
  sum(rate(http_server_requests_seconds_bucket{application="order-service"}[5m])) by (uri, le)
)
histogram_quantile(0.99, 
  sum(rate(http_server_requests_seconds_bucket{application="order-service"}[5m])) by (uri, le)
)
2. Traffic: System Load
Açıklaması şöyle
Traffic metrics include requests per second, events per second, messages per second, and batch rates. Most incidents begin with a traffic change. Sometimes expected and sometimes not.

A common pattern that I have always observed: Traffic increases, and that increases latency. Integrations slow down, and errors appear. Without traffic metrics, the root cause looks mysterious. With traffic metrics, it becomes obvious.

Prometheus query example:

Requests per second:

rate(http_server_requests_seconds_count[1m])

This metric alone explains a surprising number of incidents.
3. Errors - The most misunderstood signal
Burada şunlar önemli. 
- Error rate is more important than error count
- 4xx vs 5xx - Critical distinction

4. Saturation — Where failures actually begin
- CPU and Memory - Necessary but not enough
- Connection pool usage.
- Kubernetes saturation signals






13 Nisan 2026 Pazartesi

Production Issues Troubleshoot

Bazı problemler şöyle
Here are 15 real production scenario-based questions:

1. Your Spring Boot service CPU suddenly spikes to 90% in production. How will you investigate and fix it?

2. After deployment, your service starts throwing intermittent 500 errors. How will you debug this issue?

3. One microservice goes down and causes a chain failure in other services. How will you prevent this in future?

4. Your API response time increased from 200ms to 3 seconds after a new release. How will you identify the root cause?

5. Database connections are getting exhausted under load. What steps will you take to fix this?

6. A third-party service you depend on is timing out frequently. How will you handle this in your system?

7. You observe duplicate transactions happening in your system. How will you prevent this?

8. Logs are too large and distributed, making debugging difficult. How will you improve observability?

9. Memory usage keeps increasing and your service crashes after some time. How will you detect and fix memory leaks?

10. Your microservice works fine locally but fails in production. How will you approach debugging?

11. A new deployment breaks one feature but works for others. How will you safely roll back?

12. Traffic suddenly spikes 5x during peak hours and your service becomes slow. How will you scale?

13. Inter-service communication is failing due to network latency. How will you optimize it?

14. You need to trace a single request across multiple services during a failure. How will you implement tracing?

15. A bug in one service causes inconsistent data across multiple services. How will you handle data consistency?
Bazı problemle şöyle
Your Spring Boot service runs flawlessly in development, but crashes every night at 2am in production. Walk me through your debugging approach."

Most candidates respond:
‣ I would check the logs.
‣ I would restart the service.
‣ I would increase memory?
‣ Interview over.

Here is what interviewers are actually evaluating:

Step 1: Identify the pattern
2am is consistent. Not random. Not traffic-driven. This indicates a scheduled trigger or resource exhaustion. First question: what executes at 2am? Batch jobs? Scheduled tasks? Cron jobs?

Step 2: Analyze memory behavior before failure
Inspect JVM metrics and heap usage trends. If memory steadily increases from 10pm to 2am before crashing, it signals a memory leak not a functional bug or infrastructure issue.

Step 3: Diagnose the leak
Enable GC logs. Capture heap dumps. Identify objects with abnormal growth unclosed connections, static collections, or uncleared ThreadLocal variables. Even a single unclosed DB connection inside a loop can bring down the service.

Step 4: Validate connection pool utilization
HikariCP default pool size is 10. If a batch process consumes all connections without releasing them, subsequent requests block. By 2am, the pool is exhausted and the service becomes unresponsive.

 Solution: enforce connection timeouts and use proper try-with-resources patterns.

Step 5: Monitor with APM tools
Use Prometheus & Grafana, New Relic, or Datadog. Configure proactive alerts instead of reactive fixes. If heap usage exceeds 80% at 1am, alerts should trigger before failure occurs. That is production-grade engineering.

The gap between 12 LPA and 35 LPA is not defined by frameworks. It is defined by understanding what breaks at 3am and why.
Cpu Spike
Bir başka örnek burada

Database connections are getting exhausted under load
Örnek şöyle
@Service
public class UserService {
  @Autowired
  private JdbcTemplate jdbcTemplate; // OK

  @Transactional
  public void updateUsers(List<User> users) {
    users.forEach(user -> 
      jdbcTemplate.update(
        "UPDATE users SET last_login = ? WHERE id = ?",
        LocalDateTime.now(), user.getId()
      )
    );
 

  @Async
  @Transactional
  public void asycnUpdateUser(User user) {
    jdbcTemplate.update(
      "UPDATE users SET last_login = ? WHERE id = ?",
      LocalDateTime.now(), user.getId()
    );
  }
}
Açıklaması şöyle
Async threads can scale independently, but database connections cannot. This quickly overwhelms the connection pool.

9 Nisan 2026 Perşembe

Distributed Lock Source of Truth Olabilir mi?

Giriş
Soru şöyle
You have a distributed lock to prevent two users from booking the same hotel room.

Lock expires in 5 seconds. Your DB write takes 6 seconds under load.

Two users got confirmed bookins for the same room. How? What is the process to fix this issue.
Aslında şuna dikkat etmek lazım.
Lock ≠ correctness.
If your DB allows duplicates, your system will eventually produce them.
The real fix lives in atomic writes + constraints, not just distributed locks.
Yani lock aslında işlemi en baştan yapmamak için. Eğer iki işlem başlarsa bir tanesi başarısız olmalı.

Açıklaması şöyle
This is a correctness question. And at the Senior to Principal level, this is exactly what interviewers are testing for: do you understand the difference between coordination and actual data integrity?

If you are preparing for system design interviews right now, this is the kind of failure-mode thinking that matters a lot in strong loops.

Now, let us break this one down properly.

[1] How did both users get confirmed bookings?

The timeline usually looks like this:

- User A acquires the distributed lock for Room 101
- Lock lease is valid for 5 seconds
- User A starts the DB write to mark the room as booked
- Under load, that DB write takes 6 seconds
- At second 5, the lock expires before User A finishes
- User B now acquires the same lock because the lock service thinks it is free
- User B also starts a booking write
- Both flows eventually return success, and both users get confirmations

So what actually failed here? The system assumed the distributed lock was the source of truth. A lease-based lock only gives you temporary coordination.

If the critical section takes longer than the lease, another actor can enter while the first one is still working.

I cover fundamentals like locking, transactions, consistency, retries, idempotency, and failure handling in much more depth inside my System Design Fundamentals Guide for Senior to Principal engineers.

You can check it out here: puneetpatwari.in

[2] The deeper bug is usually not the lock itself

A lot of candidates stop at “increase the lock timeout.” That is not the real fix. The deeper issue is that your final correctness guarantee is missing at the database layer.

Because even if the lock expires, the database should still protect the invariant: “Only one valid booking can exist for this room for this date range.”

If both writes succeeded, it usually means one of these is true:
- no proper uniqueness or exclusion constraint existed
- booking availability was checked outside the final transaction
- writes were not serialized with row-level locking
- confirmation was sent before durable conflict detection finished

The lock helped reduce contention.
But the DB failed to enforce correctness.

[3] What is the right process to fix it

I would fix this in 4 steps.

1. Reconstruct the exact race
Check lock acquire time, lock expiry time, DB commit time, and confirmation event time for both users.

2. Move the invariant to the database

For hotel booking, correctness should be enforced with transactional logic such as:
- row-level locking on the inventory row
- atomic reserve-if-available update
- or exclusion/uniqueness constraints depending on data model

3. Treat the distributed lock as an optimization.
It can reduce hot contention, but it should never be the only thing preventing double booking.

4. Fix the confirmation path
Only send “booking confirmed” after the transaction commits successfully and conflict checks have passed.

5] If you still want to use distributed locks, do it safely

If a distributed lock stays in the design, I would add:
- lease renewal or heartbeats for long critical sections
- fencing tokens so stale lock holders cannot keep writing
- alerts when p99 DB latency gets too close to lock TTL
- idempotency keys so retries do not create duplicate booking flows

A good rule of thumb is simple: If your lock TTL is 5 seconds and your write path can take 6 seconds under load, your design is already telling you it is unsafe.

8 Nisan 2026 Çarşamba

Correlation Id vs Trace Id

Giriş
Açıklaması şöyle
I often noticed that some developers do not really understand the difference between traceId and correlationId. I saw this so often that I decided to write this post.

At first they look similar.
Both are IDs.
Both appear in logs.
Both help during incidents.

But they answer different questions.

traceId answers:
"How did this specific execution path go through the system?"

correlationId answers:
"Which logs and events belong to the same business story?"

That difference becomes obvious once async enters the picture 

Example:

A user places an order.

The system does this:

1. Order Service creates the order
2. Payment Service charges the card
3. Kafka event is published
4. Billing Worker creates invoice
5. Email Service sends confirmation

Now imagine the logs:

Order created
correlationId=ORDER-8472
traceId=T1

Payment charged
correlationId=ORDER-8472
traceId=T1

Billing started from Kafka consumer
correlationId=ORDER-8472
traceId=T2

Email sending failed
correlationId=ORDER-8472
traceId=T3

This is the key point 

One correlationId
Multiple traceIds

Why?

Because the business flow is one.
But the technical executions are split.

The HTTP request is one execution.
Kafka consumer is another.
Retry later can be another.
Email worker can be another too.

So:

correlationId helps you reconstruct the whole story.
traceId helps you inspect one exact path in detail.

That is why using correlationId instead of tracing is a mistake.
You may connect logs, but you still do not get spans, timing hierarchy, or where exactly latency exploded.

And using only traceId is also not enough.
In distributed async systems, tracing often shows fragments. Correlation is what lets you stitch them back together 🧩

How I usually use them during incidents:

1. Start with correlationId
Find everything related to the same order, job, or user flow.

2. Then drill into traceId
Open the exact failing execution and inspect where it slowed down or broke.

Simple version:

traceId = the path
correlationId = the story

Have you seen teams mix these two and then realize the difference only during a production incident? 

Fencing Tokens

Giriş
Açıklaması şöyle
Distributed systems concept: Fencing Tokens
You designed a fancy distributed locking algorithm just to find that an old primary is able to overwrite data!

The problem:
- Node A holds the lock, and is doing some work.
- Node A gets disconnected/unresponsive/crashes, and resume execution after its lease expires ("true" time)
- Node B, in the meantime, acquired the lock and wrote some data.
- Node A resume executions, thinking their lock is still valid
- Node A overwrites the data written by Node B, even tho it doesn't have the lock anymore.

That's were fencing token comes in: when a node acquires the lock, it gets a token with a monotonically increasing number. When the node tries to write data, it must include the token. If the token is outdated (i.e., lower than the current token), the write is rejected, preventing stale nodes from overwriting newer data.

Fencing tokens are used in a variety of systems, like etcd

The big takeaway is that you can't rely on just the client to know whether they are in their right. The target resource must have a gating mechanism to verify that the request makes sense.


JSON Web Token - JWT ve Hemen Logout

Giriş
Eğer tamamen stateless çalışıyorsak hemen logout mümkün değil. Ancak sunucu tarafına biraz state eklersek bazı çözümler elde ederiz.

1. Short-lived access tokens
- Keep access tokens valid for 5 to 15 minutes
- This limits the damage window
- Very common and simple

2. Refresh token revocation
- Store refresh tokens in DB or Redis
- On logout, delete or mark them revoked
- This is the most common real-world pattern

3. Token blacklist / denylist
- Store revoked JWT IDs or token hashes until they expire
- Check this list on every request
- Useful for high-risk logout or compromised accounts
- But now auth is no longer fully stateless

4. Token versioning
- Store a tokenVersion or sessionVersion on the user record
- Include that version in the JWT
- On logout-all-devices or password reset, increment the version
- Old tokens stop working once the version mismatches

26 Mart 2026 Perşembe

Yazılım Mimarisi - Idempotency ve Phantom Write

Giriş
Açıklaması şöyle
You typically implement idempotency like this:
  1. Check if request already processed (via key / timestamp / PK)
  2. If not → write data
  3. If yes → skip
Eğer check işlemi atomic değilse problem oluyor.

Failure Mode 1: The TTL Expiry Trap
Açıklaması şöyle
The most common idempotency implementation stores a request key with a time-to-live (TTL) — typically 24 or 48 hours. The assumption is that any duplicate will arrive within that window. In practice, this assumption frequently breaks.
Açıklaması şöyle
The fix: Never use TTL-only idempotency for operations with unbounded retry windows. Instead, use a database-backed idempotency store with a three-state model (IN_PROGRESS, COMPLETED, FAILED) where the expires_at column drives a cleanup job for storage management — not correctness. The cleanup window should be set significantly longer than your worst-case replay window (7 days minimum for Kafka-based systems).
Failure Mode 2: The Partial Execution Ghost
Açıklaması şöyle
A request arrives, the system writes the idempotency key with status IN_PROGRESS, begins processing, writes half the data, and crashes — JVM OOM, container eviction, network partition. The idempotency key is now in IN_PROGRESS state. When the retry arrives, the system faces an impossible decision: did the original operation complete or not?
Açıklaması şöyle
The fix: Wrap both the business logic and the idempotency state transition in a single database transaction. If the transaction rolls back, both the business data and the idempotency status roll back together. For stale IN_PROGRESS keys (where the original processor is likely dead), use a configurable timeout threshold to reclaim and re-execute safely.
Failure Mode 3: The Concurrent Check Race
Burada check koşulu atomic değil. Açıklaması şöyle
The fix: Use INSERT ... ON CONFLICT DO NOTHING (PostgreSQL 9.5+) to make the check-and-claim atomic. If the RETURNING clause yields no rows, the key already existed — fetch its status with SELECT ... FOR UPDATE. For non-blocking behavior, SELECT ... FOR UPDATE SKIP LOCKED lets the second instance return 409 Conflict immediately rather than waiting.
Failure Mode 4: The Layer Mismatch
Açıklaması şöyle
The fix: Propagate a correlation ID from the original request as a Kafka header, and have every downstream consumer enforce its own idempotency barrier using that ID as the deduplication key.
Spring Boot + SQL Server
Kod şöyle. Burada 
-  Partial Execution tek transaction ile çözülüyor.
- The Concurrent Check Race, DuplicateKeyException ile çözülüyor. Eğer Postgres kullanıyor olsaydık exception yerine SQL'in kaç tane satırı değiştirdiğine bakacaktır
- The Layer Mismatch sorunu outbox pattern ile çözülüyor.
@Service
@RequiredArgsConstructor
public class IdempotentService {
  private final JdbcTemplate jdbc;
  public record Response(String result) {}

  @Transactional
  public Response handleRequest(String idempotencyKey, String payload) {
    try {
      // Attempt barrier insert (atomic)
      // SQL Server:
      // INSERT INTO idempotency_table (idempotency_key, status)
      // VALUES (?, 'IN_PROGRESS')
      jdbc.update(
        "INSERT INTO idempotency_table (idempotency_key, status) VALUES (?, 'IN_PROGRESS')",
        idempotencyKey
      );

      // First request owns the key → perform business logic
      String result = doBusinessLogic(payload);

      // Insert into outbox for async processing
      // SQL Server:
      // INSERT INTO outbox_table (idempotency_key, payload) VALUES (?, ?)
      jdbc.update(
        "INSERT INTO outbox_table (idempotency_key, payload) VALUES (?, ?)",
        idempotencyKey, result
      );

      // Mark barrier as completed and store result
      // SQL Server:
      // UPDATE idempotency_table SET status='COMPLETED', response=? WHERE idempotency_key=?
      jdbc.update(
        "UPDATE idempotency_table SET status='COMPLETED', response=? WHERE idempotency_key=?",
        result, idempotencyKey
      );
      return new Response(result);
     } catch (DuplicateKeyException ex) {
      // Barrier row already exists → handle duplicate
       // SQL Server:
       // SELECT * FROM idempotency_table WITH (UPDLOCK, ROWLOCK) WHERE idempotency_key=?
       IdempotencyRecord record = jdbc.queryForObject(
         "SELECT status, response FROM idempotency_table WITH (UPDLOCK, ROWLOCK) WHERE idempotency_key=?",
         (rs, rowNum) -> new IdempotencyRecord(rs.getString("status"), rs.getString("response")),
         idempotencyKey
       );

       switch (record.status) {
         case "COMPLETED":
           // Return cached result
           return new Response(record.response);
         case "IN_PROGRESS":
           // Someone else is working → can wait or throw 409
           throw new IllegalStateException("Request is already in progress");
         case "FAILED":
           // Previous attempt failed → allow retry
           throw new IllegalStateException("Previous attempt failed, safe to retry");
         default:
           throw new IllegalStateException("Unknown barrier state: " + record.status);
         }
      }
  }

  private String doBusinessLogic(String payload) {
    // your domain logic here
    return "processed:" + payload;
  }

  private static class IdempotencyRecord {
      final String status;
      final String response;
      IdempotencyRecord(String status, String response) {
        this.status = status;
        this.response = response;
      }
  }
}
Eğer hem SQL Server hem de Postgres için çalışsın istiyorsak şöyle yaparızz
    
    
@Service
@RequiredArgsConstructor
public class IdempotentService {

    private final JdbcTemplate jdbc;

    public record Response(String result) {}

    @Transactional
    public Response handleRequest(String idempotencyKey, String payload) {
        boolean isWinner = false;

        try {
            // --------------------------
            // Attempt atomic barrier insert
            // --------------------------
            // Postgres:
            // INSERT INTO idempotency_table (idempotency_key, status)
            // VALUES (?, 'IN_PROGRESS')
            // ON CONFLICT DO NOTHING
            //
            // SQL Server:
            // INSERT INTO idempotency_table (idempotency_key, status)
            // VALUES (?, 'IN_PROGRESS')
            int rows = jdbc.update(
                    "INSERT INTO idempotency_table (idempotency_key, status) VALUES (?, 'IN_PROGRESS')",
                    idempotencyKey
            );

            // Postgres: rows == 1 → winner
            // SQL Server: INSERT succeeded → winner
            isWinner = rows == 1;

        } catch (DuplicateKeyException ex) {
            // SQL Server only: duplicate → loser
            isWinner = false;
        }

        if (isWinner) {
            // --------------------------
            // Winner executes business logic
            // --------------------------
            String result = doBusinessLogic(payload);

            // Insert into outbox (side effect)
            // INSERT INTO outbox_table (idempotency_key, payload) VALUES (?, ?)
            jdbc.update(
                    "INSERT INTO outbox_table (idempotency_key, payload) VALUES (?, ?)",
                    idempotencyKey, result
            );

            // Mark barrier as completed + store response
            // UPDATE idempotency_table SET status='COMPLETED', response=? WHERE idempotency_key=?
            jdbc.update(
                    "UPDATE idempotency_table SET status='COMPLETED', response=? WHERE idempotency_key=?",
                    result, idempotencyKey
            );

            return new Response(result);
        } else {
            // --------------------------
            // Loser reads existing row safely
            // --------------------------
            // SQL Server: SELECT ... WITH (UPDLOCK, ROWLOCK) WHERE idempotency_key=?
            // Postgres: SELECT * FROM idempotency_table WHERE idempotency_key=?
            IdempotencyRecord record = jdbc.queryForObject(
                    "SELECT status, response FROM idempotency_table " +
                            (isPostgres() ? "" : "WITH (UPDLOCK, ROWLOCK) ") +
                            "WHERE idempotency_key=?",
                    (rs, rowNum) -> new IdempotencyRecord(rs.getString("status"), rs.getString("response")),
                    idempotencyKey
            );

            switch (record.status) {
                case "COMPLETED":
                    return new Response(record.response);
                case "IN_PROGRESS":
                    throw new IllegalStateException("Request already in progress");
                case "FAILED":
                    throw new IllegalStateException("Previous attempt failed, safe to retry");
                default:
                    throw new IllegalStateException("Unknown barrier state: " + record.status);
            }
        }
    }

    private boolean isPostgres() {
        // Detect DB type from DataSource or JdbcTemplate if needed
        return true; // placeholder, implement detection
    }

    private String doBusinessLogic(String payload) {
        return "processed:" + payload;
    }

    private static class IdempotencyRecord {
        final String status;
        final String response;

        IdempotencyRecord(String status, String response) {
            this.status = status;
            this.response = response;
        }
    }
}