postgresql - postgresql mirror

	Commit message (Collapse)	Author	Age
...
*	Allow WAL record header to be split across pages.	Heikki Linnakangas	2012-06-24
\| \| \| \| \| \| \| \| \| \| \| \| \|	This saves a few bytes of WAL space, but the real motivation is to make it predictable how much WAL space a record requires, as it no longer depends on whether we need to waste the last few bytes at end of WAL page because the header doesn't fit. The total length field of WAL record, xl_tot_len, is moved to the beginning of the WAL record header, so that it is still always found on the first page where a WAL record begins. Bump WAL version number again as this is an incompatible change.
*	Move WAL continuation record information to WAL page header.	Heikki Linnakangas	2012-06-24
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The continuation record only contained one field, xl_rem_len, so it makes things simpler to just include it in the WAL page header. This wastes four bytes on pages that don't begin with a continuation from previos page, plus four bytes on every page, because of padding. The motivation of this is to make it easier to calculate how much space a WAL record needs. Before this patch, it depended on how many page boundaries the record crosses. The motivation of that, in turn, is to separate the allocation of space in the WAL from the copying of the record data to the allocated space. Keeping the calculation of space required simple helps to keep the critical section of allocating the space from WAL short. But that's not included in this patch yet. Bump WAL version number again, as this is an incompatible change.
*	Don't waste the last segment of each 4GB logical log file.	Heikki Linnakangas	2012-06-24
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The comments claimed that wasting the last segment made it easier to do calculations with XLogRecPtrs, because you don't have problems representing last-byte-position-plus-1 that way. In my experience, however, it only made things more complicated, because the there was two ways to represent the boundary at the beginning of a logical log file: logid = n+1 and xrecoff = 0, or as xlogid = n and xrecoff = 4GB - XLOG_SEG_SIZE. Some functions were picky about which representation was used. Also, use a 64-bit segment number instead of the log/seg combination, to point to a certain WAL segment. We assume that all platforms have a working 64-bit integer type nowadays. This is an incompatible change in WAL format, so bumping WAL version number.
*	Improve reporting of permission errors for array types	Peter Eisentraut	2012-06-15
\| \| \| \| \| \| \| \| \| \| \| \| \|	Because permissions are assigned to element types, not array types, complaining about permission denied on an array type would be misleading to users. So adjust the reporting to refer to the element type instead. In order not to duplicate the required logic in two dozen places, refactor the permission denied reporting for types a bit. pointed out by Yeb Havinga during the review of the type privilege feature
*	Improve readability and error messages in pg_backup_start_time.	Robert Haas	2012-06-14
\| \| \| \|	Gurjeet Singh, with corrections by me.
*	New SQL functons pg_backup_in_progress() and pg_backup_start_time()	Robert Haas	2012-06-14
\| \| \| \| \|	Darold Gilles, reviewed by Gabriele Bartolini and others, rebased by Marco Nenciarini. Stylistic cleanup and OID fixes by me.
*	During transaction cleanup, release locks before deleting files.	Robert Haas	2012-06-14
\| \| \| \| \| \| \| \|	There's no need to hold onto the locks until the files are needed, and by doing it this way, we reduce the impact on other backends who may be awaiting locks we hold. Noah Misch
*	Add new function log_newpage_buffer.	Robert Haas	2012-06-14
\| \| \| \| \| \| \| \|	When I implemented the ginbuildempty() function as part of implementing unlogged tables, I falsified the note in the header comment for log_newpage. Although we could fix that up by changing the comment, it seems cleaner to add a new function which is specifically intended to handle this case. So do that.
*	Remove RELKIND_UNCATALOGED.	Robert Haas	2012-06-14
\| \| \| \| \| \| \|	This may have been important at some point in the past, but it no longer does anything useful. Review by Tom Lane.
*	Revert "Reduce checkpoints and WAL traffic on low activity database server"	Tom Lane	2012-06-13
\| \| \| \| \| \| \| \| \| \| \| \| \|	This reverts commit 18fb9d8d21a28caddb72c7ffbdd7b96d52ff9724. Per discussion, it does not seem like a good idea to allow committed changes to go un-checkpointed indefinitely, as could happen in a low-traffic server; that makes us entirely reliant on the WAL stream with no redundancy that might aid data recovery in case of disk failure. This re-introduces the original problem of hot-standby setups generating a small continuing stream of WAL traffic even when idle, but there are other ways to address that without compromising crash recovery, so we'll revisit that issue in a future release cycle.
*	Run pgindent on 9.2 source tree in preparation for first 9.3	Bruce Momjian	2012-06-10
\| \| \| \|	commit-fest.
*	Scan the buffer pool just once, not once per fork, during relation drop.	Tom Lane	2012-06-07
\| \| \| \| \| \| \| \|	This provides a speedup of about 4X when NBuffers is large enough. There is also a useful reduction in sinval traffic, since we only do CacheInvalidateSmgr() once not once per fork. Simon Riggs, reviewed and somewhat revised by Tom Lane
*	Wake WALSender to reduce data loss at failover for async commit.	Simon Riggs	2012-06-07
\| \| \| \| \| \| \| \| \|	WALSender now woken up after each background flush by WALwriter, avoiding multi-second replication delay for an all-async commit workload. Replication delay reduced from 7s with default settings to 200ms and often much less, allowing significantly reduced data loss at failover. Andres Freund and Simon Riggs
*	Fix more crash-safe visibility map bugs, and improve comments.	Robert Haas	2012-06-07
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	In lazy_scan_heap, we could issue bogus warnings about incorrect information in the visibility map, because we checked the visibility map bit before locking the heap page, creating a race condition. Fix by rechecking the visibility map bit before we complain. Rejigger some related logic so that we rely on the possibly-outdated all_visible_according_to_vm value as little as possible. In heap_multi_insert, it's not safe to clear the visibility map bit before beginning the critical section. The visibility map is not crash-safe unless we treat clearing the bit as a critical operation. Specifically, if the transaction were to error out after we set the bit and before entering the critical section, we could end up writing the heap page to disk (with the bit cleared) and crashing before the visibility map page made it to disk. That would be bad. heap_insert has this correct, but somehow the order of operations got rearranged when heap_multi_insert was added. Also, add some more comments to visibilitymap_test, lazy_scan_heap, and IndexOnlyNext, expounding on concurrency issues. Per extensive code review by Andres Freund, and further review by Tom Lane, who also made the original report about the bogus warnings.
*	Avoid early reuse of btree pages, causing incorrect query results.	Simon Riggs	2012-06-01
\| \| \| \| \| \| \| \| \| \| \| \| \|	When we allowed read-only transactions to skip assigning XIDs we introduced the possibility that a fully deleted btree page could be reused. This broke the index link sequence which could then lead to indexscans silently returning fewer rows than would have been correct. The actual incidence of silent errors from this is thought to be very low because of the exact workload required and locking pre-conditions. Fix is to remove pages only if index page opaque->btpo.xact precedes RecentGlobalXmin. Noah Misch, reviewed by Simon Riggs
*	Improve comment for GetStableLatestTransactionId().	Tom Lane	2012-05-31
\|
*	Only throw recovery conflicts when InHotStandby. Bug fix to recent	Simon Riggs	2012-05-31
\| \| \| \| \| \|	patch to allow Index Only Scans on Hot Standby. Bug report from Jaime Casanova
*	Change the way parent pages are tracked during buffered GiST build.	Heikki Linnakangas	2012-05-30
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	We used to mimic the way a stack is constructed when descending the tree during normal GiST inserts, but that was quite complicated during a buffered build. It was also wrong: in GiST, the left-to-right relationships on different levels might not match each other, so that when you know the parent of a child page, you won't necessarily find the parent of the page to the right of the child page by following the rightlinks at the parent level. This sometimes led to "could not re-find parent" errors while building a GiST index. We now use a simple hash table to track the parent of every internal page. Whenever a page is split, and downlinks are moved from one page to another, we update the hash table accordingly. This is also better for performance than the old method, as we never need to move right to re-find the parent page, which could take a significant amount of time for buffers that were created much earlier in the index build.
*	Delete the temporary file used in buffered GiST build, after the build.	Heikki Linnakangas	2012-05-30
\| \| \| \| \| \|	There were two bugs here: We forgot to call gistFreeBuildBuffers() function at the end of build, and we passed interXact == true to BufFileCreateTemp, so the file wasn't automatically cleaned up at end-of-transaction either.
*	Fix integer overflow bug in GiST buffering build calculations.	Heikki Linnakangas	2012-05-29
\| \| \| \| \| \| \|	The result of (maintenance_work_mem * 1024) / BLCKSZ doesn't fit in a signed 32-bit integer, if maintenance_work_mem >= 2GB. Use double instead. And while we're at it, write the calculations in an easier to understand form, with the intermediary steps written out and commented.
*	Teach AbortOutOfAnyTransaction to clean up partially-started transactions.	Tom Lane	2012-05-28
\| \| \| \| \| \| \| \| \| \| \| \|	AbortOutOfAnyTransaction failed to do anything if the state it saw on entry corresponded to failing partway through StartTransaction. I fixed AbortCurrentTransaction to cope with that case way back in commit 60b2444cc3ba037630c9b940c3c9ef01b954b87b, but evidently overlooked that AbortOutOfAnyTransaction should do likewise. Back-patch to all supported branches. It's not clear that this omission has any more-than-cosmetic consequences, but it's also not clear that it doesn't, so back-patching seems the least risky choice.
*	Prevent synchronized scanning when systable_beginscan chooses a heapscan.	Tom Lane	2012-05-26
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The only interesting-for-performance case wherein we force heapscan here is when we're rebuilding the relcache init file, and the only such case that is likely to be examining a catalog big enough to be syncscanned is RelationBuildTupleDesc. But the early-exit optimization in that code gets broken if we start the scan at a random place within the catalog, so that allowing syncscan is actually a big deoptimization if pg_attribute is large (at least for the normal case where the rows for core system catalogs have never been changed since initdb). Hence, prevent syncscan here. Per my testing pursuant to complaints from Jeff Frost and Greg Sabino Mullane, though neither of them seem to have actually hit this specific problem. Back-patch to 8.3, where syncscan was introduced.
*	Ensure that seqscans check for interrupts at least once per page.	Tom Lane	2012-05-22
\| \| \| \| \| \| \| \| \| \| \| \| \|	If a seqscan encounters many consecutive pages containing only dead tuples, it can remain in the loop in heapgettup for a long time, and there was no CHECK_FOR_INTERRUPTS anywhere in that loop. This meant there were real-world situations where a query would be effectively uncancelable for long stretches. Add a check placed to occur once per page, which should be enough to provide reasonable response time without adding any measurable overhead. Report and patch by Merlin Moncure (though I tweaked it a bit). Back-patch to all supported branches.
*	Fix bug in gistRelocateBuildBuffersOnSplit().	Heikki Linnakangas	2012-05-18
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When we create a temporary copy of the old node buffer, in stack, we mustn't leak that into any of the long-lived data structures. Before this patch, when we called gistPopItupFromNodeBuffer(), it got added to the array of "loaded buffers". After gistRelocateBuildBuffersOnSplit() exits, the pointer added to the loaded buffers array points to garbage. Often that goes unnotied, because when we go through the array of loaded buffers to unload them, buffers with a NULL pageBuffer are ignored, which can often happen by accident even if the pointer points to garbage. This patch fixes that by marking the temporary copy in stack explicitly as temporary, and refrain from adding buffers marked as temporary to the array of loaded buffers. While we're at it, initialize nodeBuffer->pageBlocknum to InvalidBlockNumber and improve comments a bit. This isn't strictly necessary, but makes debugging easier.
*	Fix bug in freespace calculation in heap_multi_insert().	Heikki Linnakangas	2012-05-16
\| \| \| \| \| \| \|	If the amount of freespace on page was less than the amount reserved by fillfactor, the calculation would underflow. This fixes bug #6643 reported by Tomonari Katsumata.
*	Update comments that became out-of-date with the PGXACT struct.	Heikki Linnakangas	2012-05-14
\| \| \| \| \| \| \| \| \| \| \|	When the "hot" members of PGPROC were split off to separate PGXACT structs, many PGPROC fields referred to in comments were moved to PGXACT, but the comments were neglected in the commit. Mostly this is just a search/replace of PGPROC with PGXACT, but the way the dummy PGPROC entries are created for prepared transactions changed more, making some of the comments totally bogus. Noah Misch
*	Ensure backwards compatibility for GetStableLatestTransactionId()	Simon Riggs	2012-05-12
\|
*	Fix obsolescent C declaration syntax	Peter Eisentraut	2012-05-12
\| \| \| \| \|	gcc -Wextra/-Wold-style-declaration thinks that "inline" should go before the function return type.
*	Ensure age() returns a stable value rather than the latest value	Simon Riggs	2012-05-11
\|
*	On GiST page split, release the locks on child pages before recursing up.	Heikki Linnakangas	2012-05-11
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When inserting the downlinks for a split gist page, we used hold the locks on the child pages until the insertion into the parent - and recursively its parent if it had to be split too - were all completed. Change that so that the locks on child pages are released after the insertion in the immediate parent is done, before recursing further up the tree. This reduces the number of lwlocks that are held simultaneously. Holding many locks is bad for concurrency, and in extreme cases you can even hit the limit of 100 simultaneously held lwlocks in a backend. If you're really unlucky, you can hit the limit while in a critical section, which brings down the whole system. This fixes bug #6629 reported by Tom Forbes. Backpatch to 9.1. The page splitting code was rewritten in 9.1, and the old code did not have this problem.
*	Fix an issue in recent walwriter hibernation patch.	Tom Lane	2012-05-08
\| \| \| \| \| \| \| \| \|	Users of asynchronous-commit mode expect there to be a guaranteed maximum delay before an async commit's WAL records get flushed to disk. The original version of the walwriter hibernation patch broke that. Add an extra shared-memory flag to allow async commits to kick the walwriter out of hibernation mode, without adding any noticeable overhead in cases where no action is needed.
*	Reduce idle power consumption of walwriter and checkpointer processes.	Tom Lane	2012-05-08
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch modifies the walwriter process so that, when it has not found anything useful to do for many consecutive wakeup cycles, it extends its sleep time to reduce the server's idle power consumption. It reverts to normal as soon as it's done any successful flushes. It's still true that during any async commit, backends check for completed, unflushed pages of WAL and signal the walwriter if there are any; so that in practice the walwriter can get awakened and returned to normal operation sooner than the sleep time might suggest. Also, improve the checkpointer so that it uses a latch and a computed delay time to not wake up at all except when it has something to do, replacing a previous hardcoded 0.5 sec wakeup cycle. This also is primarily useful for reducing the server's power consumption when idle. In passing, get rid of the dedicated latch for signaling the walwriter in favor of using its procLatch, since that comports better with possible generic signal handlers using that latch. Also, fix a pre-existing bug with failure to save/restore errno in walwriter's signal handlers. Peter Geoghegan, somewhat simplified by Tom
*	Avoid repeated CLOG access from heap_hot_search_buffer.	Robert Haas	2012-05-02
\| \| \| \| \| \| \| \| \| \| \|	At the time we check whether the tuple is dead to all running transactions, we've already verified that it isn't visible to our scan, setting hint bits if appropriate. So there's no need to recheck CLOG for the all-dead test we do just a moment later. So, add HeapTupleIsSurelyDead() to test the appropriate condition under the assumption that all relevant hit bits are already set. Review by Tom Lane.
*	More duplicate word removal.	Robert Haas	2012-05-02
\|
*	Remove duplicate words in comments.	Heikki Linnakangas	2012-05-02
\| \| \| \|	Found these with grep -r "for for ".
*	Converge all SQL-level statistics timing values to float8 milliseconds.	Tom Lane	2012-04-30
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adjusts the core statistics views to match the decision already taken for pg_stat_statements, that values representing elapsed time should be represented as float8 and measured in milliseconds. By using float8, we are no longer tied to a specific maximum precision of timing data. (Internally, it's still microseconds, but we could now change that without needing changes at the SQL level.) The columns affected are pg_stat_bgwriter.checkpoint_write_time pg_stat_bgwriter.checkpoint_sync_time pg_stat_database.blk_read_time pg_stat_database.blk_write_time pg_stat_user_functions.total_time pg_stat_user_functions.self_time pg_stat_xact_user_functions.total_time pg_stat_xact_user_functions.self_time The first four of these are new in 9.2, so there is no compatibility issue from changing them. The others require a release note comment that they are now double precision (and can show a fractional part) rather than bigint as before; also their underlying statistics functions now match the column definitions, instead of returning bigint microseconds.
*	Remove duplicate word in comment.	Robert Haas	2012-04-30
\| \| \| \|	Noted by Peter Geoghegan.
*	Prevent index-only scans from returning wrong answers under Hot Standby.	Robert Haas	2012-04-26
\| \| \| \| \| \| \| \| \|	The alternative of disallowing index-only scans in HS operation was discussed, but the consensus was that it was better to treat marking a page all-visible as a recovery conflict for snapshots that could still fail to see XIDs on that page. We may in the future try to soften this, so that we simply force index scans to do heap fetches in cases where this may be an issue, rather than throwing a hard conflict.
*	Another trivial comment-typo fix.	Tom Lane	2012-04-25
\|
*	Lots of doc corrections.	Robert Haas	2012-04-23
\| \| \| \|	Josh Kupershmidt
*	Fix some typos	Peter Eisentraut	2012-04-22
\| \| \| \|	Josh Kupershmidt
*	Tighten up error recovery for fast-path locking.	Robert Haas	2012-04-18
\| \| \| \| \| \| \| \| \|	The previous code could cause a backend crash after BEGIN; SAVEPOINT a; LOCK TABLE foo (interrupted by ^C or statement timeout); ROLLBACK TO SAVEPOINT a; LOCK TABLE foo, and might have leaked strong-lock counts in other situations. Report by Zoltán Böszörményi; patch review by Jeff Davis.
*	Don't wait for the commit record to be replicated if we wrote no WAL.	Heikki Linnakangas	2012-04-17
\| \| \| \| \| \| \| \|	When using synchronous replication, we waited for the commit record to be replicated, but if we our transaction didn't write any other WAL records, that's not required because we don't even flush the WAL locally to disk in that case. This lead to long waits when committing a transaction that only modified a temporary table. Bug spotted by Thom Brown.
*	Fix typo	Peter Eisentraut	2012-04-16
\| \| \| \|	Kyotaro HORIGUCHI
*	Teach SLRU code to avoid replacing I/O-busy pages.	Robert Haas	2012-04-08
\| \| \| \|	Patch by me; review by Tom Lane and others.
*	Fix misleading output from gin_desc().	Tom Lane	2012-04-06
\| \| \| \| \| \| \| \| \| \| \|	XLOG_GIN_UPDATE_META_PAGE and XLOG_GIN_DELETE_LISTPAGE records were printed with a list link field labeled as "blkno", which was confusing, especially when the link was empty (InvalidBlockNumber). Print the metapage block number instead, since that's what's actually being updated. We could include the link values too as a separate field, but not clear it's worth the trouble. Back-patch to 8.4 where the dubious code was added.
*	Publish checkpoint timing information to pg_stat_bgwriter.	Robert Haas	2012-04-05
\| \| \| \|	Greg Smith, Peter Geoghegan, and Robert Haas
*	Correct epoch of txid_current() when executed on a Hot Standby server.	Simon Riggs	2012-03-29
\| \| \| \| \| \| \| \| \|	Initialise ckptXidEpoch from starting checkpoint and maintain the correct value as we roll forwards. This allows GetNextXidAndEpoch() to return the correct epoch when executed during recovery. Backpatch to 9.0 when the problem is first observable by a user. Bug report from Daniel Farina
*	Code cleanup for heap_freeze_tuple.	Robert Haas	2012-03-26
\| \| \| \| \| \| \|	It used to be case that lazy vacuum could call this function with only a shared lock on the buffer, but neither lazy vacuum nor any other code path does that any more. Simplify the code accordingly and clean up some related, obsolete comments.
*	Clean up compiler warnings from unused variables with asserts disabled	Peter Eisentraut	2012-03-21
\| \| \| \| \| \|	For those variables only used when asserts are enabled, use a new macro PG_USED_FOR_ASSERTS_ONLY, which expands to __attribute__((unused)) when asserts are not enabled.