postgresql - postgresql mirror

	Commit message (Collapse)	Author	Age
*	During recovery, if we reach consistent state and still have entries in the	Heikki Linnakangas	2011-12-02
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	invalid-page hash table, PANIC immediately. Immediate PANIC is much better than waiting for end-of-recovery, which is what we did before, because the end-of-recovery might not come until months later if this is a standby server. Also refrain from creating a restartpoint if there are invalid-page entries in the hash table. Restarting recovery from such a restartpoint would not see the invalid references, and wouldn't be able to cross-check them when consistency is reached. That wouldn't matter when things are going smoothly, but the more sanity checks you have the better. Fujii Masao
*	Fix getTypeIOParam to support type record[].	Tom Lane	2011-12-01
\| \| \| \| \| \| \| \| \| \| \| \| \|	Since record[] uses array_in, it needs to have its element type passed as typioparam. In HEAD and 9.1, this fix essentially reverts commit 9bc933b2125a5358722490acbc50889887bf7680, which was a hack that is no longer needed since domains don't set their typelem anymore. Before that, adjust the logic so that only domains are excluded from being treated like arrays, rather than assuming that only base types should be included. Add a regression test to demonstrate the need for this. Per report from Maxim Boguk. Back-patch to 8.4, where type record[] was added.
*	Improve table locking behavior in the face of current DDL.	Robert Haas	2011-11-30
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	In the previous coding, callers were faced with an awkward choice: look up the name, do permissions checks, and then lock the table; or look up the name, lock the table, and then do permissions checks. The first choice was wrong because the results of the name lookup and permissions checks might be out-of-date by the time the table lock was acquired, while the second allowed a user with no privileges to interfere with access to a table by users who do have privileges (e.g. if a malicious backend queues up for an AccessExclusiveLock on a table on which AccessShareLock is already held, further attempts to access the table will be blocked until the AccessExclusiveLock is obtained and the malicious backend's transaction rolls back). To fix, allow callers of RangeVarGetRelid() to pass a callback which gets executed after performing the name lookup but before acquiring the relation lock. If the name lookup is retried (because invalidation messages are received), the callback will be re-executed as well, so we get the best of both worlds. RangeVarGetRelid() is renamed to RangeVarGetRelidExtended(); callers not wishing to supply a callback can continue to invoke it as RangeVarGetRelid(), which is now a macro. Since the only one caller that uses nowait = true now passes a callback anyway, the RangeVarGetRelid() macro defaults nowait as well. The callback can also be used for supplemental locking - for example, REINDEX INDEX needs to acquire the table lock before the index lock to reduce deadlock possibilities. There's a lot more work to be done here to fix all the cases where this can be a problem, but this commit provides the general infrastructure and fixes the following specific cases: REINDEX INDEX, REINDEX TABLE, LOCK TABLE, and and DROP TABLE/INDEX/SEQUENCE/VIEW/FOREIGN TABLE. Per discussion with Noah Misch and Alvaro Herrera.
*	Tweak previous patch to ensure edata->filename always gets initialized.	Tom Lane	2011-11-30
\| \| \| \| \| \|	On a platform that isn't supplying __FILE__, previous coding would either crash or give a stale result for the filename string. Not sure how likely that is, but the original code catered for it, so let's keep doing so.
*	Strip file names reported in error messages in vpath builds	Peter Eisentraut	2011-11-30
\| \| \| \| \| \| \|	In vpath builds, the __FILE__ macro that is used in verbose error reports contains the full absolute file name, which makes the error messages excessively verbose. So keep only the base name, thus matching the behavior of non-vpath builds.
*	Prevent autovacuum transactions from running in serializable mode.	Tom Lane	2011-11-29
\| \| \| \| \| \| \| \| \| \| \| \| \|	Force the transaction isolation level to READ COMMITTED in autovacuum worker and launcher processes. There is no benefit to using a higher isolation level, and doing so could result in delaying foreground transactions (or maybe even causing unnecessary serialization failures?). Noted by Dan Ports. Also, make sure we disable zero_damaged_pages and statement_timeout in the autovac launcher, not only workers. Now that the launcher can run transactions, these settings could affect its behavior, and it seems like the same arguments apply to the launcher as the workers.
*	When a row fails a not-null constraint, show row's contents in errdetail.	Tom Lane	2011-11-29
\| \| \| \|	Simple extension of previous patch for CHECK constraints.
*	When a row fails a CHECK constraint, show row's contents in errdetail.	Tom Lane	2011-11-29
\| \| \| \| \| \| \| \| \| \| \| \|	This should make it easier to identify which row is problematic when an insert or update is processing many rows. The formatting is similar to that for unique-index violation messages, except that we limit field widths to 64 bytes since otherwise the message could get unreasonably long. (In particular, there's currently no attempt to quote or escape field values that contain commas etc.) Jan Kundrát, reviewed by Royce Ausburn, somewhat rewritten by me.
*	Make some minor formatting improvements to what pgindent did.	Tom Lane	2011-11-28
\| \| \| \| \| \|	Moving the code two full tab stops to the right requires rethinking of cosmetic code layout choices, which pgindent isn't really able to do for us. Whitespace and comment adjustments only, no code changes.
*	Disallow deletion of CurrentExtensionObject while running extension script.	Tom Lane	2011-11-28
\| \| \| \| \| \| \| \| \| \|	While the deletion in itself wouldn't break things, any further creation of objects in the script would result in dangling pg_depend entries being added by recordDependencyOnCurrentExtension(). An example from Phil Sorber convinced me that this is just barely likely enough to be worth expending a couple lines of code to defend against. The resulting error message might be confusing, but it's better than leaving corrupted catalog contents for the user to deal with.
*	Pgindent clauses.c, per request from Tom.	Bruce Momjian	2011-11-28
\|
*	Convert eval_const_expressions's long series of IsA tests into a switch.	Tom Lane	2011-11-28
\| \| \| \| \| \| \| \| \|	This function has now grown enough cases that a switch seems appropriate. This results in a measurable speed improvement on some platforms, and should certainly not hurt. The code's in need of a pgindent run now, though. Andres Freund
*	Ensure that whole-row junk Vars are always of composite type.	Tom Lane	2011-11-27
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The EvalPlanQual machinery assumes that whole-row Vars generated for the outputs of non-table RTEs will be of composite types. However, for the case where the RTE is a function call returning a scalar type, we were doing the wrong thing, as a result of sharing code with a parser case where the function's scalar output is wanted. (Or at least, that's what that case has done historically; it does seem a bit inconsistent.) To fix, extend makeWholeRowVar's API so that it can support both use-cases. This fixes Belinda Cussen's report of crashes during concurrent execution of UPDATEs involving joins to the result of UNNEST() --- in READ COMMITTED mode, we'd run the EvalPlanQual machinery after a conflicting row update commits, and it was expecting to get a HeapTuple not a scalar datum from the "wholerowN" variable referencing the function RTE. Back-patch to 9.0 where the current EvalPlanQual implementation appeared. In 9.1 and up, this patch also fixes failure to attach the correct collation to the Var generated for a scalar-result case. An example: regression=# select upper(x.*) from textcat('ab', 'cd') x; ERROR: could not determine which collation to use for upper() function
*	Use IEEE infinity, not 1e10, for null-and-not-null case in gistpenalty().	Tom Lane	2011-11-27
\| \| \| \| \| \| \|	Use of a randomly chosen large value was never exactly graceful, and now that there are penalty functions that are intentionally using infinity, it doesn't seem like a good idea for null-vs-not-null to be using something less.
*	Improve GiST range-contained-by searches by adding a flag for empty ranges.	Tom Lane	2011-11-27
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	In the original implementation, a range-contained-by search had to scan the entire index because an empty range could be lurking anywhere. Improve that by adding a flag to upper GiST entries that says whether the represented subtree contains any empty ranges. Also, make a simple mod to the penalty function to discourage empty ranges from getting pushed into subtrees without any. This needs more work, and the picksplit function should be taught about it too, but that code can be improved without causing an on-disk compatibility break; so we'll leave it for another day. Since we're breaking on-disk compatibility of range values anyway, I took the opportunity to reorganize the range flags bits; the unused RANGE_xB_NULL bits are now adjacent, which might open the door for using them in some other way later. In passing, remove the GiST range opclass entry for <>, which doesn't seem like it can really be indexed usefully. Alexander Korotkov, with some editorializing by Tom
*	Make GiST index searches smarter about queries against empty ranges.	Tom Lane	2011-11-26
\| \| \| \| \| \| \|	In the cases where the result of the called proc is negated, we should explicitly test both inputs for empty, to ensure we'll never return "true" for an unsatisfiable query. In other cases we can rely on the called proc to say the right thing.
*	Take fillfactor into account in the new COPY bulk heap insert code.	Heikki Linnakangas	2011-11-26
\| \| \| \|	Jeff Janes
*	Improve logging of autovacuum I/O activity	Alvaro Herrera	2011-11-25
\| \| \| \| \| \| \| \| \|	This adds some I/O stats to the logging of autovacuum (when the operation takes long enough that log_autovacuum_min_duration causes it to be logged), so that it is easier to tune. Notably, it adds buffer I/O counts (hits, misses, dirtied) and read and write rate. Authors: Greg Smith and Noah Misch
*	Fix erroneous replay of GIN_UPDATE_META_PAGE WAL records.	Tom Lane	2011-11-25
\| \| \| \| \| \| \| \| \| \| \| \| \|	A simple thinko in ginRedoUpdateMetapage, namely failing to increment a loop counter, led to inserting records into the last pending-list page in the wrong order (the opposite of that intended). So far as I can tell, this would not upset the code that eventually flushes pending items into the main part of the GIN index. But it did break the code that searched the pending list for matches, resulting in transient failure to find matching entries during index lookups, as illustrated in bug #6307 from Maksym Boguk. Back-patch to 8.4 where the incorrect code was introduced.
*	Move "hot" members of PGPROC into a separate PGXACT array.	Robert Haas	2011-11-25
\| \| \| \| \| \| \| \| \| \| \| \|	This speeds up snapshot-taking and reduces ProcArrayLock contention. Also, the PGPROC (and PGXACT) structures used by two-phase commit are now allocated as part of the main array, rather than in a separate array, and we keep ProcArray sorted in pointer order. These changes are intended to minimize the number of cache lines that must be pulled in to take a snapshot, and testing shows a substantial increase in performance on both read and write workloads at high concurrencies. Pavan Deolasee, Heikki Linnakangas, Robert Haas
*	Fix unsupported options in CREATE TABLE ... AS EXECUTE.	Tom Lane	2011-11-24
\| \| \| \| \| \| \| \| \| \| \|	The WITH [NO] DATA option was not supported, nor the ability to specify replacement column names; the former limitation wasn't even documented, as per recent complaint from Naoya Anzai. Fix by moving the responsibility for supporting these options into the executor. It actually takes less code this way ... catversion bump due to change in representation of IntoClause, which might affect stored rules.
*	Adjust range_adjacent to support different canonicalization rules.	Tom Lane	2011-11-23
\| \| \| \| \| \| \| \| \| \| \|	The original coding would not work for discrete ranges in which the canonicalization rule is to produce symmetric boundaries (either [] or () style), as noted by Jeff Davis. Florian Pflug pointed out that we could fix that by invoking the canonicalization function to see if the range "between" the two given ranges normalizes to empty. This implementation of Florian's idea is a tad slower than the original code, but only in the case where there actually is a canonicalization function --- if not, it's essentially the same logic as before.
*	Creator of a range type must have permission to call support functions.	Tom Lane	2011-11-23
\| \| \| \| \| \| \| \| \| \| \| \|	Since range types can be created by non-superusers, we need to consider their permissions. Ideally we'd check this when the type is used, not when it's created, but that seems like much more trouble than it's worth. The existing restriction that the support functions be immutable already prevents most cases where an unauthorized call to a function might be thought a security issue, and the fact that the user has no access to the results of the system's calls to subtype_diff closes off the other plausible reason for concern. So this check is basically pro-forma, but let's make it anyway.
*	Remove user-selectable ANALYZE option for range types.	Tom Lane	2011-11-23
\| \| \| \| \| \| \| \| \|	It's not clear that a per-datatype typanalyze function would be any more useful than a generic typanalyze for ranges. What is clear is that letting unprivileged users select typanalyze functions is a crash risk or worse. So remove the option from CREATE TYPE AS RANGE, and instead put in a generic typanalyze function for ranges. The generic function does nothing as yet, but hopefully we'll improve that before 9.2 release.
*	Remove zero- and one-argument range constructor functions.	Tom Lane	2011-11-22
\| \| \| \| \| \| \| \| \| \| \| \|	Per discussion, the zero-argument forms aren't really worth the catalog space (just write 'empty' instead). The one-argument forms have some use, but they also have a serious problem with looking too much like functional cast notation; to the point where in many real use-cases, the parser would misinterpret what was wanted. Committing this as a separate patch, with the thought that we might want to revert part or all of it if we can think of some way around the cast ambiguity.
*	Improve implementation of range-contains-element tests.	Tom Lane	2011-11-22
\| \| \| \| \| \| \| \| \| \| \| \|	Implement these tests directly instead of constructing a singleton range and then applying range-contains. This saves a range serialize/deserialize cycle as well as a couple of redundant bound-comparison steps, and adds very little code on net. Remove elem_contained_by_range from the GiST opclass: it doesn't belong there because there is no way to use it in an index clause (where the indexed column would have to be on the left). Its commutator is in the opclass, and that's what counts.
*	Check for INSERT privileges in SELECT INTO / CREATE TABLE AS.	Robert Haas	2011-11-22
\| \| \| \| \| \| \| \| \| \| \| \|	In the normal course of events, this matters only if ALTER DEFAULT PRIVILEGES has been used to revoke default INSERT permission. Whether or not the new behavior is more or less likely to be what the user wants when dealing only with the built-in privilege facilities is arguable, but it's clearly better when using a loadable module such as sepgsql that may use the hook in ExecCheckRTPerms to enforce additional permissions checks. KaiGai Kohei, reviewed by Albe Laurenz
*	Still more review for range-types patch.	Tom Lane	2011-11-22
\| \| \| \| \| \| \| \| \| \|	Per discussion, relax the range input/construction rules so that the only hard error is lower bound > upper bound. Cases where the lower bound is <= upper bound, but the range nonetheless normalizes to empty, are now permitted. Fix core dump in range_adjacent when bounds are infinite. Marginal cleanup of regression test cases, some more code commenting.
*	Continue to allow VACUUM to mark last block of index dirty	Simon Riggs	2011-11-22
\| \| \| \| \|	even when there is no work to do. Further analysis required. Revert of patch c1458cc495ff800cd176a1c2e56d8b62680d9b71
*	More code review for rangetypes patch.	Tom Lane	2011-11-21
\| \| \| \| \| \| \| \| \| \| \|	Fix up some infelicitous coding in DefineRange, and add some missing error checks. Rearrange operator strategy number assignments for GiST anyrange opclass so that they don't make such a mess of opr_sanity's table of operator names associated with different strategy numbers. Assign hopefully-temporary selectivity estimators to range operators that didn't have one --- poor as the estimates are, they're still a lot better than the default 0.5 estimate, and they'll shut up the opr_sanity test that wants to see selectivity estimators on all built-in operators.
*	Further code review for range types patch.	Tom Lane	2011-11-20
\| \| \| \| \|	Fix some bugs in coercion logic and pg_dump; more comment cleanup; minor cosmetic improvements.
*	Avoid floating-point underflow while tracking buffer allocation rate.	Tom Lane	2011-11-19
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	When the system is idle for awhile after activity, the "smoothed_alloc" state variable in BgBufferSync converges slowly to zero. With standard IEEE float arithmetic this results in several iterations with denormalized values, which causes kernel traps and annoying log messages on some poorly-designed platforms. There's no real need to track such small values of smoothed_alloc, so we can prevent the kernel traps by forcing it to zero as soon as it's too small to be interesting for our purposes. This issue is purely cosmetic, since the iterations don't happen fast enough for the kernel traps to pose any meaningful performance problem, but still it seems worth shutting up the log messages. The kernel log messages were previously reported by a number of people, but kudos to Greg Matthews for tracking down exactly where they were coming from.
*	Avoid marking buffer dirty when VACUUM has no work to do.	Simon Riggs	2011-11-18
\| \| \| \| \| \| \|	When wal_level = 'hot_standby' we touched the last page of the relation during a VACUUM, even if nothing else had happened. That would alter the LSN of the last block and set the mtime of the relation file unnecessarily. Noted by Thom Brown.
*	Further consolidation of DROP statement handling.	Robert Haas	2011-11-17
\| \| \| \| \| \| \| \| \| \| \|	This gets rid of an impressive amount of duplicative code, with only minimal behavior changes. DROP FOREIGN DATA WRAPPER now requires object ownership rather than superuser privileges, matching the documentation we already have. We also eliminate the historical warning about dropping a built-in function as unuseful. All operations are now performed in the same order for all object types handled by dropcmds.c. KaiGai Kohei, with minor revisions by me
*	Extend the unknowns-are-same-as-known-inputs type resolution heuristic.	Tom Lane	2011-11-17
\| \| \| \| \| \| \| \| \| \| \| \| \|	For a very long time, one of the parser's heuristics for resolving ambiguous operator calls has been to assume that unknown-type literals are of the same type as the other input (if it's known). However, this was only used in the first step of quickly checking for an exact-types match, and thus did not help in resolving matches that require coercion, such as matches to polymorphic operators. As we add more polymorphic operators, this becomes more of a problem. This patch adds another use of the same heuristic as a last-ditch check before failing to resolve an ambiguous operator or function call. In particular this will let us define the range inclusion operator in a less limited way (to come in a follow-on patch).
*	Fix range_cmp_bounds for the case of equal-valued exclusive bounds.	Tom Lane	2011-11-17
\| \| \| \| \| \|	Also improve its comments and related regression tests. Jeff Davis, with some further adjustments by Tom
*	Remove ancient downcasing code from procedural language operations.	Robert Haas	2011-11-17
\| \| \| \| \| \| \| \| \|	A very long time ago, language names were specified as literals rather than identifiers, so this code was added to do case-folding. But that style has ben deprecated for many years so this isn't needed any more. Language names will still be downcased when specified as unquoted identifiers, but quoted identifiers or the old style using string literals will be left as-is.
*	Restructure get_object_address() so it's safe against concurrent DDL.	Robert Haas	2011-11-17
\| \| \| \| \| \| \| \| \| \| \|	This gives a much better error message when the object of interest is concurrently dropped and avoids needlessly failing when the object of interest is concurrently dropped and recreated. It also improves the behavior of two concurrent DROP IF EXISTS operations targeted at the same object; as before, one will drop the object, but now the other will emit the usual NOTICE indicating that the object does not exist, instead of rolling back. As a fringe benefit, it's also slightly less code.
*	Improve caching in range type I/O functions.	Tom Lane	2011-11-15
\| \| \| \| \|	Cache the the element type's I/O info across calls, not only the range type's info. In passing, also clean up hash_range a bit more.
*	Restructure function-internal caching in the range type code.	Tom Lane	2011-11-15
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Move the responsibility for caching specialized information about range types into the type cache, so that the catalog lookups only have to occur once per session. Rearrange APIs a bit so that fn_extra caching is actually effective in the GiST support code. (Use of OidFunctionCallN is bad enough for performance in itself, but it also prevents the function from exploiting fn_extra caching.) The range I/O functions are still not very bright about caching repeated lookups, but that seems like material for a separate patch. Also, avoid unnecessary use of memcpy to fetch/store the range type OID and flags, and don't use the full range_deserialize machinery when all we need to see is the flags value. Also fix API error in range_gist_penalty --- it was failing to set *penalty for any case involving an empty range.
*	Fix alignment and toasting bugs in range types.	Tom Lane	2011-11-14
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	A range type whose element type has 'd' alignment must have 'd' alignment itself, else there is no guarantee that the element value can be used in-place. (Because range_deserialize uses att_align_pointer which forcibly aligns the given pointer, violations of this rule did not lead to SIGBUS but rather to garbage data being extracted, as in one of the added regression test cases.) Also, you can't put a toast pointer inside a range datum, since the referenced value could disappear with the range datum still present. For consistency with the handling of arrays and records, I also forced decompression of in-line-compressed bound values. It would work to store them as-is, but our policy is to avoid situations that might result in double compression. Add assorted regression tests for this, and bump catversion because of fixes to built-in pg_type entries. Also some marginal cleanup of inconsistent/unnecessary error checks.
*	Return NULL instead of throwing error when desired bound is not available.	Tom Lane	2011-11-14
\| \| \| \| \| \| \| \|	Change range_lower and range_upper to return NULL rather than throwing an error when the input range is empty or the relevant bound is infinite. Per discussion, throwing an error seems likely to be unduly hard to work with. Also, this is more consistent with the behavior of the constructors, which treat NULL as meaning an infinite bound.
*	Return FALSE instead of throwing error for comparisons with empty ranges.	Tom Lane	2011-11-14
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Change range_before, range_after, range_adjacent to return false rather than throwing an error when one or both input ranges are empty. The original definition is unnecessarily difficult to use, and also can result in undesirable planner failures since the planner could try to compare an empty range to something else while deriving statistical estimates. (This was, in fact, the cause of repeatable regression test failures on buildfarm member jaguar, as well as intermittent failures elsewhere.) Also tweak rangetypes regression test to not drop all the objects it creates, so that the final state of the regression database contains some rangetype objects for pg_dump testing.
*	Fix copyright notices, other minor editing in new range-types code.	Tom Lane	2011-11-14
\| \| \| \| \| \| \|	No functional changes in this commit (except I could not resist the temptation to re-word a couple of error messages). This is just manual cleanup after pgindent to make the code look reasonably like other PG code, in preparation for more detailed code review to come.
*	Rerun pgindent with updated typedef list.	Bruce Momjian	2011-11-14
\|
*	Run pgindent on range type files, per request from Tom.	Bruce Momjian	2011-11-14
\|
*	Wakeup WALWriter as needed for asynchronous commit performance.	Simon Riggs	2011-11-13
\| \| \| \| \| \| \| \|	Previously we waited for wal_writer_delay before flushing WAL. Now we also wake WALWriter as soon as a WAL buffer page has filled. Significant effect observed on performance of asynchronous commits by Robert Haas, attributed to the ability to set hint bits on tuples earlier and so reducing contention caused by clog lookups.
*	Avoid retaining multiple relation locks in RangeVarGetRelid.	Robert Haas	2011-11-12
\| \| \| \| \| \| \| \|	If it turns out we've locked the wrong OID, release the old lock. In most cases, it's pretty harmless to retain the extra lock, but this seems tidier and avoids using lock table slots unnecessarily. Per discussion with Tom Lane.
*	Revert removal of trace_userlocks, because userlocks aren't gone.	Robert Haas	2011-11-10
\| \| \| \| \| \|	This reverts commit 0180bd6180511875db046bf8ddcaa633a2952dfd. contrib/userlock is gone, but user-level locking still exists, and is exposed via the pg_advisory* family of functions.
*	Fix another bug in the redo of COPY batches.	Heikki Linnakangas	2011-11-10
\| \| \| \| \|	I got alignment wrong in the redo routine. Spotted by redoing the log genereated by copy regression test.