Prev Next

Integration / Apache Pulsar Interview questions

How can you optimize BookKeeper storage costs using tiered storage?

Configure an offload threshold (by size or age) at the namespace or topic level so ledger segments older than the threshold are automatically moved from replicated BookKeeper disk to cheaper object storage like S3, without manual intervention once set up.

Choose the threshold deliberately based on actual access patterns: frequently read "hot" data should stay on BookKeeper's low-latency disks, while data rarely accessed except for occasional replay or compliance should be pushed to offload as early as reasonably possible to minimize the replicated-disk footprint.

Because offloaded segments remain transparently readable through the normal topic/consumer APIs, you can safely set aggressive (short) offload thresholds without breaking any existing consumer or replay workflow - the only cost is somewhat higher read latency the rare times older data actually gets accessed.

Combine this with a sensible retention policy: for data you don't need indefinitely, letting it expire and delete outright, rather than retaining and offloading forever, is cheaper still than paying even discounted object storage costs for data nobody will ever read again.

Tiered storage reduces cost mainly by moving old data:
Offloaded data being transparently readable means you can:

Invest now in Acorns!!! 🚀 Join Acorns and get your $5 bonus!
Acorns Logo

Invest now in Acorns!!! 🚀
Join Acorns and get your $5 bonus!

Earn passively and while sleeping

Acorns is a micro-investing app that automatically invests your "spare change" from daily purchases into diversified, expert-built portfolios of ETFs. It is designed for beginners, allowing you to start investing with as little as $5. The service automates saving and investing. Disclosure: I may receive a referral bonus.

Robinhood Logo

Invest now!!! Get Free equity stock (US, UK only)!

Use Robinhood app to invest in stocks. It is safe and secure. Use the Referral link to claim your free stock when you sign up!.

The Robinhood app makes it easy to trade stocks, crypto and more.


Webull Logo

Webull! Receive free stock by signing up using the link: Webull signup.

More Related questions...

Explain about Apache Pulsar. What is Apache Pulsar? What is a Pulsar broker? What is Apache BookKeeper? What are tenants and namespaces in Pulsar? What is a topic in Pulsar? What is a producer in Pulsar? What is a consumer in Pulsar? What is a subscription in Pulsar? Define a partitioned topic in Pulsar? What is Pulsar's metadata store used for? What are Pulsar Functions? What is Pulsar IO? Describe geo-replication in Pulsar? What is tiered storage in Pulsar? What is a non-persistent topic in Pulsar? What are the subscription types in Pulsar? What is message retention in Pulsar? What is schema registry in Pulsar? List the core components of a Pulsar cluster? How do you create a topic in Pulsar? What is the difference between Pulsar and Kafka's storage architecture? How does Pulsar separate compute and storage? Why is Pulsar considered multi-tenant by design? What is the difference between Shared and Exclusive subscriptions? How does Key_Shared subscription maintain ordering? When should you use Failover subscription instead of Exclusive? What is the difference between a ledger and a segment in BookKeeper? How does Pulsar achieve message deduplication? Why do brokers in Pulsar not store data locally? What happens when a broker crashes in Pulsar? How does namespace bundle splitting work? What is the difference between backlog quota and retention policy? When should you use a Reader instead of a Consumer? How does topic compaction work in Pulsar? Why is ensemble size different from write quorum in BookKeeper? What is the difference between persistent and non-persistent topics? How does Pulsar handle delayed message delivery? What happens when a consumer negatively acknowledges a message? Explain the lifecycle of a message in Pulsar from produce to acknowledge? How can you optimize Pulsar for high-throughput workloads? How do you troubleshoot a growing backlog in Pulsar? Explain the execution flow of topic ownership failover in Pulsar? How can you optimize BookKeeper storage costs using tiered storage? Explain the internal working of Pulsar transactions? Which is better for exactly-once processing: idempotent producers or transactions, and why? How do you troubleshoot unbalanced load across brokers? Explain the lifecycle of a namespace bundle from creation to split? How can you optimize consumer throughput with Key_Shared subscriptions? Explain the execution flow of a Pulsar Function processing a message? How do you troubleshoot message duplication in a Pulsar producer?
Show more question and Answers...

Apache Camel Interview Questions

Comments & Discussions